Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Research Scientist, Model Evaluation”. A match may be a passing mention rather than the job itself. Titles only.
201 roles across 212 listings · show every listing · page 1 of 9
…As a Senior Research Scientist, Model Evaluation, you will: - Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish…
…ABOUT THE ROLE We're looking for the first dedicated Data Scientist to partner with the Identity organization. In this role, you will define…
…We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0…
…Senior Applied Scientist in the Alexa AI team: - Define and drive the science roadmap for conversational AI capabilities powered by large language models - Design…
…Key job responsibilities - Lead the research, design, and development of advanced AI/ML systems for information retrieval and agentic systems - Identify and evaluate emerging…
…Previous experience with the ML lifecycle (data collection and labeling, model evaluation, etc.). Applied experience using data to inform priorities and evaluate impact. Excellent…
…of scientists and engineers, fostering a collaborative and inclusive culture of continuous learning Establish and enforce best practices for model monitoring, evaluation, and performance…
…Key job responsibilities • Collaborate with AI/ML scientists and architects to research, design, develop, and evaluate generative AI solutions to address real-world challenges…
…As a Senior Scientist for the team, you will have the opportunity to apply your deep subject matter expertise in the area of ML…
…e.g., foundation models for control, diffusion policies, world models) and assess their applicability to safe legged locomotion • Publish research at top-tier robotics…
…end ownership of ML models — from data collection and labeling strategy to training, evaluation, and deployment • Mentor junior scientists and engineers; contribute to a…
…strategy, influence candidate evaluation, and guide hiring decisions that directly impact the direction and quality of our frontier-model research and fulfillment of our…
…Partner closely with product managers, engineers, data scientists, and designers to define and execute experimentation strategies. Drive A/B testing, monitoring, model evaluation, and…
…equivalent practical experience.8+ years of senior technical leadership experience influencing engineers, technical leads, architects, applied scientists, or cross-functional engineering teams across complex…
…Your morning might start with reviewing model performance metrics and experiment results, collaborating with Applied Scientists to optimize LLM prompting strategies or model architectures…
…We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0…
…the Role We are seeking a Senior Software Engineer with deep expertise in scheduling and operations research to join the Autonomous Lab software organization…
…research scientist, you'll be at the center of that work. You'll drive a high-impact research agenda focused on large language models…
…tool surfaces, orchestration, prompt pipelines, evaluation harnesses, and the backend APIs that make AI-enabled workflows safe and observable. They also deliver full-stack…
…Senior Software Engineer, Autonomous Lab About the Role We are seeking Senior Software Engineers to join the Autonomous Lab software organization at Ginkgo Bioworks…
…of Applied Scientists, Research Scientists, and Machine Learning Engineers. - Architect and scale multi-agent systems - Partner with Product, Engineering, and senior leadership (including S…
…As a Senior Applied Scientist, you will develop and improve machine learning systems that enable real-time manufacturing flow decisions. You will leverage state…
…We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0…
…We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0…
…Our work spans AI evaluation and benchmarking, data development, model improvement, and operational AI deployment in some of the most demanding environments in government…