Site Reliability Engineer, Model Post Training
Thinking Machines LabAICore Engineering
Under a week · posted
Half of Thinking Machines Lab’s 43 open roles have been sitting longer than 22 days.
Thinking Machines Lab posted this Site Reliability Engineer, Model Post Training role in San Francisco; New York today, 26 August 2026. fable12 indexed it directly from Thinking Machines Lab's Ashby board.
What Thinking Machines Lab says about it
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. ABOUT THE ROLE We're looking for a Site Reliability Engineer (SRE) to drive reliability for Model Post Training end-to-end. Post-training spans several distinct workloads — supervised fine-tuning, reinforcement learning, preference optimization, and distillation — each with its own resource profile, failure signature, and tolerance for interruption or delay. You'll set the reliability strategy across all of them, not just react to what breaks. This is a senior individual-contributor role with real technical authority.
The opening of the posting. Read it on Thinking Machines Lab’s own board →
Email me when Thinking Machines Lab posts another role → One a day at most, nothing on a quiet day.
Details
- Location
- San FranciscoNew York
- Workplace
- On site
- Employment
- Full time
- Team
- Core Engineering
- Source
- Thinking Machines Lab’s Ashby boardLast seen 27 August 2026