Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Scientist - RL Training”. A match may be a passing mention rather than the job itself. Titles only.
14 roles across 15 listings · show every listing
…Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family…
…This in-office expectation does not apply to contractor positions About the role We are looking for a founding senior Research Scientist with a…
…You will work closely with research scientists and product engineers on multimodal data processing, model training, inference and serving tasks . As a contributing member…
…researchers with the technical depth to move the frontier in pretraining, reinforcement learning (RL) / post-training, or evals; and researchers with real depth in…
…Experience with reinforcement learning (RL) and post-training with LLMs, especially spoken LLMs. Excellent engineering skills in Python and deep learning frameworks (e.g…
…As a Research Scientist, you'll also actively contribute to the wider research community by sharing and publishing your findings, with ideas inspired by…
…Job Description What you'll do Own LLM mid-training and post-training research, including continued pretraining, SFT, preference optimization, and RL; make data…
…from collection and curation through to preprocessing, quality assurance, and delivery into training pipelines. You'll work closely with research scientists to understand what…
…We are looking for AI research scientists, Applied scientists and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion…
…RL agents) in both cloud environments and air-gapped, offline tactical edge networks. You will be a force multiplier for our AI Research Scientists…
…agent orchestration frameworks) - Have experience with reward modeling, RLHF/RLAIF/RLVR, or preference-based training - Have industrial or academic experience with CAE or EDA…
…Deep technical grounding across multiple areas of applied AI, including LLM post-training (RLHF, DPO, distillation), model routing and adaptation, agentic systems, code generation…
…and post-training methodologies for large language models specialized in code intelligence - Invent and implement state-of-the-art agentic RL training recipes that…
…Training, RL & Evaluation Infrastructure • Build and scale the offline inference systems behind post-training — high-throughput rollout generation and reward-model serving for reinforcement…