Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Research Scientist (Measurement and Evaluation)”. A match may be a passing mention rather than the job itself. Titles only.
469 roles across 522 listings · show every listing · page 2 of 19
…reproducible, scalable, and well-documented analytical workflows Evaluate model performance through backtesting, forecast-versus-actual reporting, and clearly defined accuracy measures Required Qualifications Build…
…forecasting, and model evaluation; fluent in SQL and Python. - Proven ability to turn messy, broad datasets into high-impact decisions and measurable outcomes. - Enthusiasm…
…Operating with minimal oversight, you will translate clinical and product questions into rigorous research approaches, generate and interpret evidence from clinical, wearable, and real…
…right way to measure success has to be invented rather than applied. - Run tight feedback loops into Product, Engineering, and Research, shaping the product…
…As a Research Scientist, you'll setup large-scale tests and deploy promising ideas quickly and broadly, managing deadlines and deliverables while applying the…
…context, or evaluation. Independently identify research problems, formulate novel hypotheses, and drive projects from research idea to validated prototype and measurable impact, collaborating with…
…Design and conduct A/B tests to evaluate algorithm performance and propose iterative improvements based on the results. Effectively communicate findings, insights, and recommendations…
…You'll partner directly with applied scientists on measuring adherence and effectiveness of agentic tooling and with product managers to translate ambiguous customer problems…
…Decision Making Design and evaluate business interventions using experimentation, quasi-experimentation, and causal inference. Build clear analytical frameworks and measurement approaches that improve decision…
…Design and own evaluation frameworks (unit evals, integration evals, production monitors) that measure agent quality, catch regressions, and drive data-informed decisions LLM integration…
…learning, and AI evaluation, partnering closely with software engineers, researchers, and product managers to develop data-driven approaches that measure, analyze, and improve AI…
…Success requires strong technical judgment, partnership with legal and regulatory experts, disciplined evaluation, human-in-the-loop design, clear communication, and a commitment to…
…We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0…
…We are looking for AI research scientists, Applied scientists and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion…
…Optimize AI Researcher Velocity: Actively identify, measure, and eliminate bottlenecks in the ML research lifecycle. Build highly automated tools for hyperparameter tuning, model profiling…
…actually work, and with the broader research team to translate that domain grounding into training signal and evaluation benchmarks that measure genuine task competence…
…Conduct deep market research, competitive analysis, and customer interviews to uncover untapped opportunities and validate core product hypotheses. MVP & Validation: Design rapid experiments (prototypes…
…Build solution patterns, evaluation frameworks, and playbooks. Extract what works across engagements and feed field signals back to Product and Research to improve our…
…and acceptance criteria Design hypotheses and run experiments (such as A/B tests) to validate impact and measure outcomes Partner with data scientists and…
…You will establish scientific and technical direction, lead research and applied-development efforts, influence architecture and product strategy, and mentor other scientists and engineers…
…Mentor senior engineers and applied scientists on building reliable, scalable, and reusable evaluation infrastructure. Stay current with LLM evaluation methods, agentic systems, RAG evaluation…
…you can reason about evaluation, retrieval quality, tool use, failure modes, and what is/isn’t worth building. - Experience applying evaluation and observability to…
…that span both AI research scientists and production engineers, including hiring, developing talent, and making difficult personnel decisions. Strong evaluation fluency, with a track…
…Define success metrics, experimentation frameworks (A/B, causal inference), and measurement methodology for seller engagement interventions. Productionize ML models and data products — partner with…
…and measure impact), and operations (to ground models in how the network actually runs), and use modern GenAI/LLM tooling to accelerate research and…