Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Software Engineer, RL Post-Training Frameworks”. A match may be a passing mention rather than the job itself. Titles only.
10 roles across 11 listings · show every listing
…Work at the intersection of compter-architecture, libraries, frameworks, AI applications and the entire software stack. Innovate and improve model architectures, distributed training algorithms…
…Expertise in post-training of large language models, including: Policy optimization algorithms (e.g. GRPO, PPO) Verifier-based RL frameworks (RLVR) Supervised fine-tuning…
…Expertise in post-training of large language models, including: Policy optimization algorithms (e.g. GRPO, PPO) Verifier-based RL frameworks (RLVR) Supervised fine-tuning…
SENIOR SOFTWARE ENGINEER, MACHINE LEARNING INFRASTRUCTURE - GENERATIVE AI ABOUT THE TEAM Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the…
…Grow and steward the expert contributor network of professional software engineers across languages, frameworks, and seniority levels, balancing throughput, cost, and quality. Own and…
…Analyze complex hardware-software interactions to identify and resolve performance bottlenecks in both training and inference pipelines. Collaborate closely with AI researchers, HW and…
…This includes scaling ladders for pre-training, data mix optimizations and RL recipes (preferences) for post-training. Engage with the wider research community on…
…you'll own discrete pieces of real systems under the mentorship of senior engineers, not shadow work or isolated coursework-style projects. What You…
…There is also appetite on this team for post-training our own models where an internal workload justifies it, and the engineer in this…
…In this hybrid role, you will report to a Senior Staff Software Engineer. You will: Partner with foundation model training teams to determine optimal…