Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Software Engineer, RL Post-Training Frameworks”. A match may be a passing mention rather than the job itself. Titles only.
42 roles across 50 listings · show every listing · page 1 of 2
…NVIDIA is building an RL Frameworks engineering team to develop the open-source tools and infrastructure that AI researchers and post-training teams depend…
…and Trainium Systems — delivering high-performance ML inference and training at cloud scale. We’re looking for a Senior Manager, Physical Design Engineering to…
…data generation, and preference data collection for RLHF/DPO-style training. Own data quality: build validation frameworks, monitor for drift and contamination, and establish…
…Training, RL & Evaluation Infrastructure • Build and scale the offline inference systems behind post-training — high-throughput rollout generation and reward-model serving for reinforcement…
…Work at the intersection of compter-architecture, libraries, frameworks, AI applications and the entire software stack. Innovate and improve model architectures, distributed training algorithms…
…Expertise in post-training of large language models, including: Policy optimization algorithms (e.g. GRPO, PPO) Verifier-based RL frameworks (RLVR) Supervised fine-tuning…
…Expertise in post-training of large language models, including: Policy optimization algorithms (e.g. GRPO, PPO) Verifier-based RL frameworks (RLVR) Supervised fine-tuning…
SENIOR SOFTWARE ENGINEER, MACHINE LEARNING INFRASTRUCTURE - GENERATIVE AI ABOUT THE TEAM Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the…
…Grow and steward the expert contributor network of professional software engineers across languages, frameworks, and seniority levels, balancing throughput, cost, and quality. Own and…
…Analyze complex hardware-software interactions to identify and resolve performance bottlenecks in both training and inference pipelines. Collaborate closely with AI researchers, HW and…
…This includes scaling ladders for pre-training, data mix optimizations and RL recipes (preferences) for post-training. Engage with the wider research community on…
…you'll own discrete pieces of real systems under the mentorship of senior engineers, not shadow work or isolated coursework-style projects. What You…
…There is also appetite on this team for post-training our own models where an internal workload justifies it, and the engineer in this…
…In this hybrid role, you will report to a Senior Staff Software Engineer. You will: Partner with foundation model training teams to determine optimal…
…batching, KV cache optimization, RLHF, RLAIF, DPO, or other post-training and alignment methods. Experience creating workload qualification frameworks, benchmark plans, or technical decision…
…pre-training, production RL runs, evaluation, and agentic AI infrastructure. We build the software foundations that help research and engineering teams train, evaluate, and…
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of…
…As a Senior Machine Learning Engineer, you will own the ML lifecycle for the language models that understand and reason about the content in…
…To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine…
…Strong engineering and R&D experience in LLM post-training, reinforcement learning for language models (RLHF, RLAIF), reward modeling, policy optimization, model alignment, grounding…
…Fine-tuning (PEFT, SFT), post-training and RL from verifiable rewards, Reasoning, RAG, Agent Evaluations and Observability, and Production inference. Represent partner needs and…
…ABOUT THE ROLE We’re looking for a Senior Software Engineer to join our ML Infrastructure & Platform team. This team powers both Handshake’s…
…Experience with model customization or post-training techniques such as SFT, RL/RLHF/RLAIF, DPO or relevant equivalent experience, reward modeling, LoRA/PEFT, quantization…
…sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use…
…platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of…