Skip to content
fable12

Site Reliability Engineer, Model Post Training

Thinking Machines LabAICore Engineering

under a daysince it was posted

Under a week · posted

Half of Thinking Machines Lab’s 43 open roles have been sitting longer than 22 days.

Apply at Thinking Machines LabOpens Thinking Machines Lab’s own Ashby board

Thinking Machines Lab posted this Site Reliability Engineer, Model Post Training role in San Francisco; New York today, 26 August 2026. fable12 indexed it directly from Thinking Machines Lab's Ashby board.

What Thinking Machines Lab says about it

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. ABOUT THE ROLE We're looking for a Site Reliability Engineer (SRE) to drive reliability for Model Post Training end-to-end. Post-training spans several distinct workloads — supervised fine-tuning, reinforcement learning, preference optimization, and distillation — each with its own resource profile, failure signature, and tolerance for interruption or delay. You'll set the reliability strategy across all of them, not just react to what breaks. This is a senior individual-contributor role with real technical authority.

The opening of the posting. Read it on Thinking Machines Lab’s own board →

Email me when Thinking Machines Lab posts another role → One a day at most, nothing on a quiet day.

Details

Location
San FranciscoNew York
Workplace
On site
Employment
Full time
Team
Core Engineering
Source
Thinking Machines Lab’s Ashby boardLast seen 27 August 2026

Also open at Thinking Machines Lab

All 43