Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Engineer, Code RL (Reinforcement Learning)”. A match may be a passing mention rather than the job itself. Titles only.
12 roles
…engineering, research, and everything in between. Your contributions will span model architecture, data curation, training and inference infrastructures, evaluation protocols, alignment and reinforcement learning…
…researchers with the technical depth to move the frontier in pretraining, reinforcement learning (RL) / post-training, or evals; and researchers with real depth in…
…ABOUT THE ROLE Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts…
…Experience with reinforcement learning (RL) and post-training with LLMs, especially spoken LLMs. Excellent engineering skills in Python and deep learning frameworks (e.g…
…Science, Machine Learning, Robotics, Control Engineering, or a closely related field Possess solid software engineering skills, writing clean and well-structured code in Python…
…Drive post-training research and engineering using reinforcement learning (RL) and supervised fine-tuning (SFT) to advance Gemini coding capabilities across web, 3D, game…
…engineering automated evaluation harnesses, reward models, or testing frameworks (e.g., LLM-as-a-judge, visual regression testing, Reinforcement Learning from Human Feedback (RLHF…
…agentic AI (planning, tool use, memory), reinforcement learning (RLHF, RLVF, RLGF, offline RL). One or more scientific publication submissions for conferences, journals, or public…
…Hands-on experience implementing Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) or direct preference optimization (DPO) loops. Multimodal Experience: Experience working with multimodal…
…FoSSE is an applied research team focused on solving real-world software engineering (SE) problems for large codebases, including code migration, code translation, and…
…in agentic RL space: - Push the boundaries of reinforcement learning and post-training methodologies for large language models specialized in code intelligence - Invent and…
…reward-model serving for reinforcement learning (RL/RLHF/RLAIF) • Ensure train/serve consistency — that the inference path used in RL and evaluation faithfully matches…