Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Engineer, Performance RL (Reinforcement Learning)”. A match may be a passing mention rather than the job itself. Titles only.
14 roles across 15 listings · show every listing
…dexterous manipulation and relevant frontier research problems - Conduct original research on world-action foundation models and reinforcement learning, delivering manipulation policies for humanoid robotic…
…Experience with reinforcement learning (RL) and post-training with LLMs, especially spoken LLMs. Excellent engineering skills in Python and deep learning frameworks (e.g…
…GPU-based simulator to train reinforcement-learning agents across thousands of parallel environments Develop the tooling that underpins RL training, including domain randomisation, to…
…Drive post-training research and engineering using reinforcement learning (RL) and supervised fine-tuning (SFT) to advance Gemini coding capabilities across web, 3D, game…
…engineering automated evaluation harnesses, reward models, or testing frameworks (e.g., LLM-as-a-judge, visual regression testing, Reinforcement Learning from Human Feedback (RLHF…
…agentic AI (planning, tool use, memory), reinforcement learning (RLHF, RLVF, RLGF, offline RL). One or more scientific publication submissions for conferences, journals, or public…
…Hands-on experience implementing Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) or direct preference optimization (DPO) loops. Multimodal Experience: Experience working with multimodal…
…reinforcement learning alignment loops (RLHF/DPO) to guarantee model safety and predictability in high-stakes environments. Collaborate and Pioneer: Work closely with AI Researchers…
…API orchestration, Planning, large multimodal models (especially vision-language models), reinforcement learning (RL) and sequential decision making. * Define and implement new automated reasoning features…
…API orchestration, Planning, large multimodal models (especially vision-language models), reinforcement learning (RL) and sequential decision making. * Define and implement new automated reasoning features…
…API orchestration, Planning, large multimodal models (especially vision-language models), reinforcement learning (RL) and sequential decision making. About the team Agentic AI drives innovation…
…API orchestration, Planning, large multimodal models (especially vision-language models), reinforcement learning (RL) and sequential decision making. About the team Agentic AI drives innovation…
…We build many of these reinforcement learning (RL) environments, then drop our agents into them to evaluate or train them. In this role, you…
…reward-model serving for reinforcement learning (RL/RLHF/RLAIF) • Ensure train/serve consistency — that the inference path used in RL and evaluation faithfully matches…