Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer - Automated Evaluation and Adversarial Design”. A match may be a passing mention rather than the job itself. Titles only.
13 roles across 14 listings · show every listing
…Key job responsibilities - Design, develop, and deploy production ML systems for abuse pattern detection, anomaly detection, threat classification, and automated enforcement across multiple Amazon…
…Design and build end-to-end automated evaluation pipelines that streamline the validation workflow, from input ingestion and response evaluation to scoring and reporting…
…across automated detections, human review, appeals, chargebacks, and support cases, and establish a trustworthy baseline for the first time. - Design and evaluate a risk…
…design, and adoption across engineering teams. - Hands-on experience evaluating or implementing AI security tooling, including AI-augmented testing, LLM security assessments, or automated…
…case-management platforms, or decision engines to automate complex operational workflows. A strong understanding of LLM-agent evaluation and reliability, including hallucination control, grounding…
…that evaluate every aspect of our systems Develop automated testing frameworks to enable continuous assessment at scale Collaborate with Product, Engineering, and Policy teams…
…Integrate LLM tool use, function calling, retrieval, planning loops, evaluation hooks, and guardrails with simulation engines, physics/modeling backends, and mission planning systems. Build…
…engineering action - Strong fundamentals in cryptography, identity/access management, and secure software development lifecycle NICE TO HAVE - Experience securing GPU infrastructure or ML training…
…MLOps & Automation: Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI…
…Proven RL system design in live environments (optimization/control/fraud decisioning), including reward design, online/offline evaluation, and safe deployment in adversarial settings. Experience…
…Experience designing agent evaluation paradigms, including trajectory evaluations, LLM-as-judge workflows, task-success metrics, tool-call correctness checks, rubric-based qualitative grading, adversarial…
…Partner with Product Platform, Cloud Infrastructure, and Data engineering teams to ensure core primitives, APIs, and microservices are secure by default from design to…
…Policy, Security, Identity, and partner engineering teams to translate business needs into scalable platform capabilities. Drive productionization of AI/ML and agentic capabilities with…