Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Engineer, Code RL (Reinforcement Learning)”. A match may be a passing mention rather than the job itself. Titles only.
119 roles · group by role · page 1 of 5
…Research Engineer, Performance RL (Reinforcement Learning) — teaching Claude to write correct, fast code for accelerators Research Engineer, Universes — long-horizon, ultra-realistic agentic training…
…engineering, research, and everything in between. Your contributions will span model architecture, data curation, training and inference infrastructures, evaluation protocols, alignment and reinforcement learning…
…researchers with the technical depth to move the frontier in pretraining, reinforcement learning (RL) / post-training, or evals; and researchers with real depth in…
…ABOUT THE ROLE Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts…
…Experience with reinforcement learning (RL) and post-training with LLMs, especially spoken LLMs. Excellent engineering skills in Python and deep learning frameworks (e.g…
…Science, Machine Learning, Robotics, Control Engineering, or a closely related field Possess solid software engineering skills, writing clean and well-structured code in Python…
…Drive post-training research and engineering using reinforcement learning (RL) and supervised fine-tuning (SFT) to advance Gemini coding capabilities across web, 3D, game…
…engineering automated evaluation harnesses, reward models, or testing frameworks (e.g., LLM-as-a-judge, visual regression testing, Reinforcement Learning from Human Feedback (RLHF…
…agentic AI (planning, tool use, memory), reinforcement learning (RLHF, RLVF, RLGF, offline RL). One or more scientific publication submissions for conferences, journals, or public…
…Hands-on experience implementing Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) or direct preference optimization (DPO) loops. Multimodal Experience: Experience working with multimodal…
…FoSSE is an applied research team focused on solving real-world software engineering (SE) problems for large codebases, including code migration, code translation, and…
…in agentic RL space: - Push the boundaries of reinforcement learning and post-training methodologies for large language models specialized in code intelligence - Invent and…
…reward-model serving for reinforcement learning (RL/RLHF/RLAIF) • Ensure train/serve consistency — that the inference path used in RL and evaluation faithfully matches…
…vetting design, solution approaches, results and code. Deep research expertise in multi-agent systems and reinforcement learning including agent co-ordination, negotiation, communication, simulations…
…Define and maintain clear research quality standards and engineering best practices and engineering standards for the team PhD degree in Computer Science, Machine Learning…
…OUR IDEAL STAFF RESEARCH SCIENTIST, EXOTIC AI WILL HAVE: - 8+ years of relevant experience in machine learning engineering, AI research, or a closely related…
…ML codebases and distributed systems. - Experience improving model behavior through data, reward modeling, or RL techniques. - Evidence of owning ambitious research or engineering agendas…
…high-fidelity simulations where AI learns to perform real-world tasks through reinforcement learning. We work with the leading AI labs to help them…
…We build training gyms for AI agents — high-fidelity simulations where models learn to solve economically valuable problems through reinforcement learning. We work with…
…high-fidelity simulations where AI learns to perform real-world tasks through reinforcement learning. We work with the leading AI labs to help them…
…high-fidelity simulations where AI learns to perform real-world tasks through reinforcement learning. We work with the leading AI labs to help them…
…You’ll gain hands on experience with real data, production infrastructure and real deadlines, while learning from and working alongside the researchers and engineerings…
…You’ll gain hands on experience with real data, production infrastructure and real deadlines, while learning from and working alongside the researchers and engineerings…
…of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Team Our Reinforcement Learning teams are…
…reinforcement learning and related learning-based control for high-DOF, multi-fingered robotic hands—translating state-of-the-art research (reinforcement / imitation learning, teleoperation…