Skip to content

Open role · AI Quality & Evaluation

AI Evaluation Engineer

Build the datasets, harnesses, failure taxonomies, and decision evidence that separate promising AI behavior from production-ready performance.

Remote — internationalRemoteFull-time

Job description

About the role

Evaluation is one of Innomium’s core engineering disciplines. This role turns ambiguous claims—“the agent works,” “the model understands the document,” “the detector is accurate”—into protocols a delivery team can inspect and act on.

The mandate

You will design evaluations for LLM, agent, retrieval, vision, and applied AI systems. The work includes task definition, representative test cases, annotation guidance, automated and human scoring, adversarial cases, slice analysis, regression infrastructure, and communication of uncertainty.

You will work with researchers and engineers before implementation is complete, helping define what success means and which evidence should stop or redirect the work. You will also help teams avoid metric theatre: optimizing a convenient average while the operational failure modes remain hidden.

What strong performance looks like

You create an evaluation system that is repeatable, decision-oriented, and maintainable. Results are segmented by meaningful conditions, failures are translated into engineering work, and regressions become visible before they reach users.

How we work

This role requires technical depth and editorial clarity. You should be able to work in code, challenge a research claim constructively, facilitate expert review, and explain why a result is or is not sufficient for the next delivery gate.

Responsibilities

The work this role is expected to own.

  • Translate product and operating outcomes into measurable AI acceptance criteria
  • Create representative datasets, annotation guidance, rubrics, and failure-mode taxonomies
  • Build automated evaluation and regression harnesses for LLM, agent, retrieval, or vision systems
  • Design human-review workflows and reconcile subjective or expert judgments
  • Analyze results by meaningful slices and turn failure patterns into engineering priorities
  • Maintain evaluation lineage, versioning, reproducibility, and decision reports

Requirements

Capabilities and experience that support success in this role.

  • Professional experience evaluating machine-learning or AI-enabled products
  • Strong Python and data-analysis skills with the ability to build maintainable tooling
  • Understanding of statistical uncertainty, dataset construction, bias, and measurement error
  • Ability to define rubrics and communicate nuanced results without overstating certainty
  • Experience collaborating with researchers, engineers, domain experts, or product teams
  • Evidence of rigorous analytical work and clear technical writing

Nice to have

Useful adjacent experience, but not a substitute for the core requirements.

  • Experience with LLM-as-judge design, human preference evaluation, or red teaming
  • Experience with computer-vision evaluation or production monitoring
  • Knowledge of regulated, safety-sensitive, or high-stakes decision environments

How to apply

Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.

Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.

Email your application

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.