Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Scientist - RL Training”. A match may be a passing mention rather than the job itself. Titles only.
23 roles across 25 listings · show every listing
…YOUR RESPONSIBILITIES We are looking for a Senior Research Scientist to lead fine-tuning, post-training, and reinforcement learning for the next generation of…
…YOUR RESPONSIBILITIES We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next…
…prototyping, ablations, scaling experiments, evaluation, and delivery into production, with rigorous and reproducible evaluation. - Partner closely with post-training, RL/RLHF, and instruction-following…
…looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal…
…research techniques. Focus on managing a team of scientists that would work on scaling ladders for pre-training, data mix optimizations and RL recipes…
…bring scientific innovations into production. - Stay current with the latest research in LLMs, RL, and agent-based AI, optimization and translate findings into practical…
…leveraging latest advances in RL-based fine-tuning methods like DPO, GRPO etc. - Identifying latest technical/research trends applicable for the problems of efficient…
…in collaboration with RL researchers to train LLMs capable of directing complex simulation pipelines. - Generate simulated datasets for ML training in regimes where experimental…
…Strong engineering and R&D experience in LLM post-training, reinforcement learning for language models (RLHF, RLAIF), reward modeling, policy optimization, model alignment, grounding…
…Meta's Research & Development teams are looking for a Research Scientist who will invent and productize advanced AI model–hardware codesign techniques. Working at…
…5+ years of experience as an ML engineer or applied/research scientist, including direct experience training or fine-tuning models in production systems PhD…
…post-training and RL environments, scaled post-training infrastructure, and strategic project support across our lab partnerships. You think like an applied researcher, lead…
…training and neural network architecture development, visualization of complex multi-modal sensor data, and more. You will collaborate with other engineers and researchers across…
…Research engineers and scientists with experience building and training computer vision models. Experience with multimodal representations and visual language modeling is strongly preferred. A…
…Redis, or feature platform infrastructure. - Experience with post-training techniques such as fine-tuning, RLHF, reinforcement learning, or reward modeling. - Experience building agentic systems…
…Post-Training — Research in post-training techniques for music generation, including preference alignment methods (such as DPO, RLHF, or KTO), reward model design and…
…of LLM/agent research and Canva's product goals. Model Innovation: Creating novel agent capabilities through post-training and RL—developing reward modeling, synthetic…
…SAN FRANCISCO, CA | HYBRID, 3X A WEEK IN OFFICE WHAT YOU'LL DO: - Partner directly with AI lab researchers to understand their post-training…
…Turing accelerates frontier research with high-quality data, specialized talent, and training pipelines that advance thinking, reasoning, coding, multimodality, and STEM. For enterprises, Turing…
…Collaborate with Data Scientists, Researchers, and Engineers to drive improvements across our platforms. We are looking for an Evaluation & Insights Engineer for the Human…
…domain scientists. What matters is strong AI and robotics foundations, scientific curiosity, and the drive to ship. Key job responsibilities - Develop, train, and benchmark…
As an Applied Scientist II in the Alexa Conversational Modelling Intelligence team within Alexa AI, you will drive model post-training for Large Language…
…scientists and engineers at Amazon developing the next generation of safe autonomy, while also establishing strong connections with top academic research labs. Your research…