Skip to content

Open role · AI Research

Research Engineer, Long-Context Models

Investigate long-context architectures, training behavior, kernels, and evaluation while turning research findings into reproducible engineering artifacts.

Remote — internationalRemoteFull-time

Job description

About the role

Innomium’s long-context research includes Continuum1-9B and its supporting linear-attention stack. This role sits between model research and systems engineering: experiments must be technically ambitious, reproducible, and relevant to a decision about what should be built next.

The mandate

You will investigate linear and hybrid attention, recurrent state-update mechanisms, distillation, long-sequence training, information retention, and the execution path required to make experiments practical. You may work across model architecture, data mixtures, training instrumentation, custom kernels, and evaluation.

The best candidate is comfortable reading papers and kernels, but does not confuse novelty with value. You will compare long-context approaches with retrieval, hierarchy, and simpler baselines, and you will document where evidence is incomplete.

What strong performance looks like

You can define a falsifiable research question, build a controlled experiment, diagnose unexpected behavior, and publish an artifact or decision record another engineer can reproduce. Over time, your work improves the quality and efficiency of Innomium’s model and systems research.

How we work

Research is reviewed through hypotheses, experiment records, ablations, and limitations. You will collaborate with evaluation, infrastructure, and product engineers so that model advances remain connected to realistic workloads and operating constraints.

Responsibilities

The work this role is expected to own.

  • Design and run controlled experiments on long-context and linear-attention architectures
  • Develop or adapt training pipelines, model code, data curricula, and instrumentation
  • Investigate optimization, stability, state behavior, extrapolation, and long-sequence failure modes
  • Collaborate on Triton or CUDA kernel integration and hardware-level performance analysis
  • Build reproducible evaluation paths spanning generic benchmarks and long-context stress tests
  • Write clear research notes, model documentation, experiment records, and limitation statements

Requirements

Capabilities and experience that support success in this role.

  • Strong foundation in deep learning and modern language-model architectures
  • Advanced Python and PyTorch experience with distributed training or model internals
  • Ability to translate papers into controlled implementations and critical experiments
  • Understanding of attention, optimization, numerical behavior, and evaluation methodology
  • Evidence of research engineering through repositories, model releases, papers, or substantial experiments
  • Careful technical writing and willingness to report negative or ambiguous results

Nice to have

Useful adjacent experience, but not a substitute for the core requirements.

  • Experience with linear attention, state-space models, distillation, or extreme context
  • Triton, CUDA, FlashAttention, or performance-kernel experience
  • Experience publishing open-weight models or Hugging Face custom modeling code

How to apply

Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.

Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.

Email your application

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.