Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Software Engineer, RL Post-Training Frameworks”. A match may be a passing mention rather than the job itself. Titles only.
16 roles across 18 listings · show every listing
…NVIDIA is building an RL Frameworks engineering team to develop the open-source tools and infrastructure that AI researchers and post-training teams depend…
…batching, KV cache optimization, RLHF, RLAIF, DPO, or other post-training and alignment methods. Experience creating workload qualification frameworks, benchmark plans, or technical decision…
…pre-training, production RL runs, evaluation, and agentic AI infrastructure. We build the software foundations that help research and engineering teams train, evaluate, and…
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of…
…As a Senior Machine Learning Engineer, you will own the ML lifecycle for the language models that understand and reason about the content in…
…To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine…
…Strong engineering and R&D experience in LLM post-training, reinforcement learning for language models (RLHF, RLAIF), reward modeling, policy optimization, model alignment, grounding…
…Fine-tuning (PEFT, SFT), post-training and RL from verifiable rewards, Reasoning, RAG, Agent Evaluations and Observability, and Production inference. Represent partner needs and…
…ABOUT THE ROLE We’re looking for a Senior Software Engineer to join our ML Infrastructure & Platform team. This team powers both Handshake’s…
…Experience with model customization or post-training techniques such as SFT, RL/RLHF/RLAIF, DPO or relevant equivalent experience, reward modeling, LoRA/PEFT, quantization…
…sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use…
…platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of…
…optimization libraries, or distributed computing frameworks with public upstream records. Solid experience in Agentic AI, RL post-training or long-context LLM workload optimization…
…or defect/quality analysis. - Familiarity with modern training/inference infrastructure (e.g., distributed training, RL frameworks, model serving). Amazon is an equal opportunities employer…
…scale distributed AI training, post-training (RLHF, alignment), inference, or robotics workloads (multi-node, multi-GPU) Experience driving framework or software bring-up on…
…Experience building internal ML platforms or research clusters at a company doing large-scale training Familiarity with agentic AI: RL training with rollouts, agent…