Site Reliability Engineer, RL Infra
Thinking Machines LabAICore Engineering
Under a week · posted · closed
Half of Thinking Machines Lab’s 38 open roles have been sitting longer than 23 days.
Thinking Machines Lab kept this Site Reliability Engineer, RL Infra listing in San Francisco; New York open for under a day, from 26 August 2026. It closed on 27 August 2026. fable12 indexed it directly from Thinking Machines Lab's Ashby board.
What Thinking Machines Lab says about it
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. ABOUT THE ROLE We're looking for a Site Reliability Engineer (SRE) to drive reliability for RL Infra end-to-end. RL training is unlike a typical batch training job: it interleaves rollout generation, environment or tool interactions, reward scoring, and policy weight updates in a continuous loop, often across long-tailed trajectories with unpredictable latency. Keeping this loop healthy — and keeping it healthy for many concurrent tenants sharing the same underlying clusters — is the core of this role.
The opening of the posting. Read it on Thinking Machines Lab’s own board →
Details
- Location
- San FranciscoNew York
- Workplace
- On site
- Employment
- Full time
- Team
- Core Engineering
- Source
- Thinking Machines Lab’s Ashby boardLast seen 27 August 2026