Skip to content

← Back to all roles

Engineering · Remote (US)

ML Engineer, Post-Training

Own evals, fine-tuning, and the production behavior of the agents we deploy: routing, drafting, extraction, and the judgment boundary between model and human.

About the role

The agents Zoro installs draft replies, route money, extract records, and decide what a human must approve. When one of them is wrong, a real company feels it. Post-training is where that behavior gets owned.

You will own the eval harnesses, the fine-tuning runs, and the routing layer that decides which model sees which task. The hardest part of the job is the judgment boundary: what the model may do alone, and what it must hand to a person.

What you will do

  • Own the eval suites that gate every model change before production
  • Fine-tune and route models for drafting, extraction, and routing tasks
  • Define and enforce the judgment boundary between model and human
  • Watch production behavior and drive failure rates down week over week
  • Lower cost per correct action without lowering the standard

What we are looking for

  • LLM systems shipped to production, not notebooks and demos
  • Deep eval methodology: you can say exactly why a model got better
  • Fine-tuning experience: SFT, preference tuning, or both
  • Strong Python and strong opinions about measurement
  • You read papers and ship code, in that order

Nice to have

  • Open-weight model deployments under real latency and cost limits
  • Retrieval systems over messy company data
  • Observability tooling for agent behavior in production