Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer - Automated Evaluation and Adversarial Design”. A match may be a passing mention rather than the job itself. Titles only.
67 roles across 72 listings · show every listing · page 1 of 3
…This role focuses on building and scaling automated evaluation systems and designing adversarial and stress-testing methodologies across multiple AI features. The work requires…
…Design and execute A/B tests and counterfactual analyses to evaluate the efficacy and side effects of anti-spam rules, ML model updates, and…
…systems engineering, infrastructure, security, networking, data and analytics - Knowledge of distributed systems design and implementation or equivalent - Knowledge of large scale automation and workflow…
…Hands-on experience with AI technologies, whether through building ML models, working with LLMs and prompt engineering, experimenting with agentic frameworks, or applying AI…
…and post-launch evaluation to ensure fraud-awareness by design Ecosystem Monitoring & Knowledge Leadership - Continuously survey external fraud trends, adversary techniques, tooling, and emerging…
…Scale E2E ML systems. Collaborate with engineering on data contracts, feature stores, distributed training/inference, and automated rollout/rollback; drive architectural investments that increase…
…Develop multimodal evaluation pipelines reasoning across text, images, audio, and video in real time - Design adaptive ML systems resilient to adversarial evolution-- continuously learning…
…Design secure, high-quality AI and software architectures, reviewing and challenging designs and code to ensure adversarial resilience. Reduce AI and LLM security vulnerabilities…
…Design secure, high-quality AI and software architectures, reviewing and challenging designs and code to ensure adversarial resilience. Reduce AI and LLM security vulnerabilities…
…MLOps & Automation: Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI…
…You'll build and manage a small team of engineers, decide how to extend that team's reach through automation and strategic partners, and…
…Security engineers on your team translate threat intelligence into the adversary behaviors that matter; you build and tune the models that detect those behaviors…
…consequences - - Experience designing observability and instrumentation systems for production ML or AI workloads, including trace collection, evaluation harnesses, and cost and latency monitoring for…
…and consumer contracts while mentoring junior engineers on technical quality and design practices. MS/PhD in a quantitative field (e.g., Statistics, ML) or…
…productizing the execution runtime and hardening the security boundary around it; driving generation quality through prompt iteration, automated evaluation, and output validation; designing the…
…Security engineers on your team translate threat intelligence into the adversary behaviors that matter; you build the models that detect those behaviors and evaluate…
We are looking for a Senior Solutions Architect to help leading Enterprise ISVs design, build, and deploy secure agentic AI systems on NVIDIA’s…
…You will work in a highly collaborative and cross-functional environment, partnering with ML Engineers, Trust & Safety Ops, subject matter expert teams, and Product…
…into automated systems, and prioritize high-impact risk areas * Partner with engineering teams to deploy models into production, define evaluation frameworks, and collaborate with…
…Evaluate end-to-end behavior through automated scoring, trajectory analysis, and adversarial testing. Use evaluation results, production traces, and analyst feedback to improve agent…
…Design secure, high-quality AI and software architectures, reviewing and challenging designs and code to ensure adversarial resilience, secure-by-default patterns, and appropriate…
…defenses, and red-teaming or adversarial evaluation practices Hands-on experience with observability and evaluation tools for LLMs (e.g., LangSmith, Weights & Biases, MLflow)
…LoRA/QLoRA, instruction tuning, RLHF/DPO, dataset curation, and evaluating tuned vs prompted performance. Python + ML engineering proficiency: PyTorch, Hugging Face, LangChain/LlamaIndex (or…
…of the ML and data stack (e.g., pandas, scikit-learn, PyTorch or TensorFlow) for feature engineering, model training, and automation. Solid understanding of…
…and over-linkage risk. - Define evaluation methods for Digital Intelligence signals, including holdout design, leakage checks, drift monitoring, adversarial robustness, customer impact analysis, and…