Skip to content
fable12

Site Reliability Engineer, RL Infra

Thinking Machines LabAICore Engineering

under a daybefore it closed

Under a week · posted · closed

Half of Thinking Machines Lab’s 38 open roles have been sitting longer than 23 days.

View the original postingLeft Ashby — the page may be gone

Thinking Machines Lab kept this Site Reliability Engineer, RL Infra listing in San Francisco; New York open for under a day, from 26 August 2026. It closed on 27 August 2026. fable12 indexed it directly from Thinking Machines Lab's Ashby board.

What Thinking Machines Lab says about it

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. ABOUT THE ROLE We're looking for a Site Reliability Engineer (SRE) to drive reliability for RL Infra end-to-end. RL training is unlike a typical batch training job: it interleaves rollout generation, environment or tool interactions, reward scoring, and policy weight updates in a continuous loop, often across long-tailed trajectories with unpredictable latency. Keeping this loop healthy — and keeping it healthy for many concurrent tenants sharing the same underlying clusters — is the core of this role.

The opening of the posting. Read it on Thinking Machines Lab’s own board →

Details

Location
San FranciscoNew York
Workplace
On site
Employment
Full time
Team
Core Engineering
Source
Thinking Machines Lab’s Ashby boardLast seen 27 August 2026

Also open at Thinking Machines Lab

All 38