Job description
About the role
Evaluation is one of Innomium’s core engineering disciplines. This role turns ambiguous claims—“the agent works,” “the model understands the document,” “the detector is accurate”—into protocols a delivery team can inspect and act on.
The mandate
You will design evaluations for LLM, agent, retrieval, vision, and applied AI systems. The work includes task definition, representative test cases, annotation guidance, automated and human scoring, adversarial cases, slice analysis, regression infrastructure, and communication of uncertainty.
You will work with researchers and engineers before implementation is complete, helping define what success means and which evidence should stop or redirect the work. You will also help teams avoid metric theatre: optimizing a convenient average while the operational failure modes remain hidden.
What strong performance looks like
You create an evaluation system that is repeatable, decision-oriented, and maintainable. Results are segmented by meaningful conditions, failures are translated into engineering work, and regressions become visible before they reach users.
How we work
This role requires technical depth and editorial clarity. You should be able to work in code, challenge a research claim constructively, facilitate expert review, and explain why a result is or is not sufficient for the next delivery gate.
Responsibilities
The work this role is expected to own.
- Translate product and operating outcomes into measurable AI acceptance criteria
- Create representative datasets, annotation guidance, rubrics, and failure-mode taxonomies
- Build automated evaluation and regression harnesses for LLM, agent, retrieval, or vision systems
- Design human-review workflows and reconcile subjective or expert judgments
- Analyze results by meaningful slices and turn failure patterns into engineering priorities
- Maintain evaluation lineage, versioning, reproducibility, and decision reports
Requirements
Capabilities and experience that support success in this role.
- Professional experience evaluating machine-learning or AI-enabled products
- Strong Python and data-analysis skills with the ability to build maintainable tooling
- Understanding of statistical uncertainty, dataset construction, bias, and measurement error
- Ability to define rubrics and communicate nuanced results without overstating certainty
- Experience collaborating with researchers, engineers, domain experts, or product teams
- Evidence of rigorous analytical work and clear technical writing
Nice to have
Useful adjacent experience, but not a substitute for the core requirements.
- Experience with LLM-as-judge design, human preference evaluation, or red teaming
- Experience with computer-vision evaluation or production monitoring
- Knowledge of regulated, safety-sensitive, or high-stakes decision environments
How to apply
Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.
Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.
Email your application