Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Engineer, Code RL (Reinforcement Learning)”. A match may be a passing mention rather than the job itself. Titles only.
114 roles across 119 listings · show every listing · page 4 of 5
…fine-tuning, reinforcement learning, and on-policy distillation. Experience developing, evaluating, and deploying machine learning models in production environments. Strong research and problem solving…
…other engineers. Preferred Qualifications: Publications, open source contributions, or other demonstrated research impact in reinforcement learning, LLM post-training, or agent learning. Deep experience…
…supports reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops. - Work closely with research teams to support distributed RL workloads…
…Developing systems that enable models to use computers effectively Advancing code generation through reinforcement learning Pioneering fundamental RL research for large language models Building…
…engineering, research, and everything in between. Your contributions will span model architecture, data curation, training and inference infrastructure, evaluation protocols, alignment and reinforcement learning…
…from reinforcement learning and upstream training through to deployment of standalone, customer-facing products. The ideal candidate is equal parts researcher, engineer, and product…
…Reinforcement learning training infrastructure - Distributed training and inference systems - Experiment infrastructure and reproducibility - Large-scale data pipelines The goal is to build the engineering…
…Agentic systems using deep reinforcement learning to solve hard problems. Where necessary, you will also work on integrating ML/RL frameworks into our products…
…to read and translate research papers and prototypes into shippable engineering Strong candidates also have - Experience with reinforcement learning - Background in building evaluation frameworks…
…and consumer applications The Staff Research Engineer – Code will own end-to-end the creation of datasets, RL environments, and evals for frontier AI…
…You will work closely with researchers and senior engineers to implement and improve workflows for LLM pretraining, fine-tuning, and reinforcement learning-based post…
…a Machine Learning Engineer or Research Scientist, applying research to tangible products. Deep technical understanding of Imitation Learning, Robotics, Reinforcement Learning, Computer Graphics and…
…Developing systems that enable models to use computers effectively Advancing code generation through reinforcement learning Pioneering fundamental RL research for large language models Building…
…Experience fine-tuning or adapting foundation models using methods like supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and low-rank adaptation…
…We are looking for engineers who can navigate the convergence of machine learning and systems engineering to build robust, scalable platforms. What you will…
…Implement environment abstractions to support reinforcement learning and agent evaluation at scale. Collaborate on large scale RL-infra: RL-trainer, rollout system, and containerized…
…Use reinforcement learning (policy optimization, bandits, RLHF‑style approaches where appropriate) to improve personalization, dialog policies, and sequential decision‑making systems. Fine-tune and…
…Strong software engineering capabilities with experience building automated evaluation pipelines or large-scale ML systems. - Experience with Reinforcement Learning (RLHF/RLAIF) and how it…
…Working at the intersection of machine learning and systems engineering, you'll integrate reinforcement learning agents into high-performance runtime systems, enabling real-time…
…Supervised Finetuning (SFT), Reinforcement Learning (RL), prompt improvements and synthetic data generation Work in concert with product and infrastructure engineers to improve Figma’s…
…following areas: - Post-training and reinforcement learning: Techniques used to improve model deployment quality through further training, tuning, RL, and focus on particular downstream…
…About Horizons The Horizons team leads Anthropic's reinforcement learning (RL) research and development, playing a critical role in advancing our AI systems. We…
…Familiarity with RL-specific infrastructure requirements (e.g., actor/learner architectures, experience replay systems, large-scale environment execution). Strong software engineering practices: code quality…
…engineering, research, and everything in between. Your contributions will span model architecture, data curation, training and inference infrastructures, evaluation protocols, alignment and reinforcement learning…
…engineering, research, and everything in between. Your contributions will span model architecture, data curation, training and inference infrastructure, evaluation protocols, alignment and reinforcement learning…