KarminaAI StudioAgents · Services · Training

JOBS · Collaboration by project

AI Evals & Reliability Engineer

It designs tests to understand when an agent is reliable, when it fails, and how it should improve.

The opportunity

It “seems to work” in evidence.

We are looking for a profile of AI Evaluat & Reliability Engineer to design the testing layer of agents and generative applications. We want to measure behavior, quality, safety, cost and latency with representative cases before relying on a demo or a subjective impression.

Area
Quality and reliability
Experience
3+ years
Modality
Remote
Collaborating
By Project or Recurring

What are we looking for?

  • Experience creating test datasets, acceptance criteria and regression suites.
  • Knowledge of deterministic assessment, similarity, classification, model-based graders and human review.
  • Ability to analyze traces, tool calls, recoveries and final results.
  • Experience with Python, data analysis and test automation.
  • Criteria for distinguishing useful metrics from scores that do not reflect the actual process.
  • Ability to communicate failures and trade-offs in an understandable way.

What are you gonna do?

  • Turn business requirements and risks into test cases.
  • Build banks of correct, incorrect, ambiguous and adversarial examples.
  • Design response, retrieval, tool use, handoffs and boundary compliance evals.
  • Analyze regressions when changing model, prompt, sources, or tools.
  • Measure cost, latency, escalation rate, false positives and human review effort.
  • Prepare dashboards and reports to decide if a pilot advances, corrects or stops.
  • Coordinate testing with domain specialists and process managers.

It will impress us especially.

  • A suite that has discovered a major mistake before production.
  • Experience with network teaming, LLM observability, evaluation of RAG or agent testing with tools.
  • Know when an automatic metric needs a well-designed human sample.

What will you find?

The evaluation will not be a final phase: it will participate from the blueprint and will accompany each relevant change. You will have the authority to show that an apparent improvement does not exceed the set of tests.

How to introduce yourself

Attach your CV and, if you want, a single link to a GitHub, an evaluation framework, a test plan or an anonymized dashboard.

Talent

AI Evals & Reliability Engineer

This is the way to submit a candidacy or propose a professional collaboration. All fields are optional. Project and service inquiries are managed at hello@karmina.ai.

jobs@agenciakarmina.com