Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Member of Technical Staff - Post-Training and RL”. A match may be a passing mention rather than the job itself. Titles only.
44 roles · page 1 of 2
…You are obsessed about building incredibly useful models through post-training and RL techniques. You are a power user of AI models and eager…
…The foundational models require large compute-capacity, and as a Member of Technical Staff – Multimodal Infrastructure you would be responsible to build large-scale…
…Develop, pre-train, and fine-tune in-house LLMs and multimodal foundation models. Apply SOTA post-training alignment techniques (SFT, RLHF, DPO) to maximize…
…As a Member of Technical Staff, RL Environments, you will: - Build new RL environments targeting different agentic capabilities and industry areas - Train and evaluate…
…Have experience with reward modeling, RL, or other post-training techniques. Preferred Qualifications Bachelor's Degree in Computer Science or related technical field AND…
…powerful pre-trained models into aligned and general agents. - Drive research and engineering initiatives that push the frontier of post-training, from data curation…
…Understanding of ML/LLM concepts. Post-training, evaluation metrics, reward modeling: deep enough to partner with researchers and execute on novel RL environments. - Thrives…
…the life - Train ML models for deployment in simulation and real-world robots, identify and document their limitations post-deployment - Drive technical discussions within…
…for training and inference, sized to support frontier-scale experimentation including large-model pre-training and post-training, RL training runs, and large-scale…
…Deep proficiency in Python, and comfort across the rest of the stack. As an RL post-training practitioner You have fine-tuned models for…
…post-training modern language or multimodal models. - Strong understanding of machine learning fundamentals and current post-training and RL methods. - Solid engineering skills and…
…RL-based post-training (PPO/DPO/GRPO), and large-scale distributed training (data, model, and context parallelism). - A history of being a technical authority…
…Experience with custom model training, fine-tuning, or post-training (SFT, RLHF/DPO) over proprietary technical data. Excellent self-motivation, creativity, and a passion…
…training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and…
…training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and…
…Optimization and integration of inference systems into our RL training stack. CORE TECHNICAL RESPONSIBILITIES LLM Serving - Multi‑tenant LLM Serving: Build a multi-tenant…
…training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and…
…training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and…
…training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and…
…proprietary models, training data, and the compute that powers them. This role owns the security posture of everything we ship: the hosted RL training…
…technical pillars include post-training enhancements, harness design, agentic reinforcement learning (RL), and the construction of specialized RL environments. As a member of this…
…engineers who have strong judgment and set technical direction, quickly build prototypes that scale into the reliable systems, and are at the frontier of…
…in technical discussions about new model architectures with the science team - Manage pre/post training runs and continue improve system stability and throughput - Prototype…
…This is a Staff / Senior IC role. We're looking for someone who has shipped post-training for a frontier model before and wants…
…Work across the stack (pre-training → SFT/RL/post-training) to enable reasoning, tool calling, agentic behaviors, orchestration, and seamless real-time interactions. BASIC…