Engineering · Remote (US)
ML Engineer, Post-Training
Own evals, fine-tuning, and the production behavior of the agents we deploy: routing, drafting, extraction, and the judgment boundary between model and human.
About the role
The agents Zoro installs draft replies, route money, extract records, and decide what a human must approve. When one of them is wrong, a real company feels it. Post-training is where that behavior gets owned.
You will own the eval harnesses, the fine-tuning runs, and the routing layer that decides which model sees which task. The hardest part of the job is the judgment boundary: what the model may do alone, and what it must hand to a person.
What you will do
- Own the eval suites that gate every model change before production
- Fine-tune and route models for drafting, extraction, and routing tasks
- Define and enforce the judgment boundary between model and human
- Watch production behavior and drive failure rates down week over week
- Lower cost per correct action without lowering the standard
What we are looking for
- LLM systems shipped to production, not notebooks and demos
- Deep eval methodology: you can say exactly why a model got better
- Fine-tuning experience: SFT, preference tuning, or both
- Strong Python and strong opinions about measurement
- You read papers and ship code, in that order
Nice to have
- Open-weight model deployments under real latency and cost limits
- Retrieval systems over messy company data
- Observability tooling for agent behavior in production