Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer, Inference & Optimization”. A match may be a passing mention rather than the job itself. Titles only.
1,291 roles across 1,495 listings · show every listing · page 3 of 52
…Key job responsibilities Define, develop and deploy innovative physical design and verification methodologies (RTL2GDS) for ML Accelerator chips in advanced nodes Drive Optimizations in…
…ML Engineer/Data Engineer/Data Scientist, Experience with machine learning techniques and advanced analytics (e.g. regression, classification, clustering, time series, econometrics, causal inference…
…On-device ML experience (model size optimization, quantization, power-efficient inference) is a strong plus. Background in signal processing, Kalman filtering, particle filters, or…
…you can connect methods to business decisions, not just optimize offline metrics. The ability to operate across functions with ML engineers, economists, data scientists…
…you can connect methods to business decisions, not just optimize offline metrics. The ability to operate across functions with ML engineers, economists, data scientists…
…You can engage deeply with scientists on evaluation methodology, data quality, and model behavior, and equally deeply with engineers on architecture, inference latency, and…
…Optimize and deploy models into production autonomous-driving systems, working across model architecture, inference, and onboard constraints. Set technical direction for geometric vision at…
…You'll collaborate across engineering, product, and business teams to uncover opportunities, optimize experiences for organizations, and influence decisions at global scale. This role…
…in Electrical Engineering, Computer Engineering, Computer Science, or equivalent practical experience 12+ years architecting hardware systems for hyperscale, HPC, or AI/ML infrastructure Deep…
…or Kubernetes for ML job orchestration Experience with hyperparameter optimization and experiment tracking tools\ Background in ML Engineering, AI Engineering, MLOps, or LLMOps Prior…
…Collaborate cross-functionally with product managers, researchers, and engineers to deliver secure, high-quality, and scalable AI/ML systems. What you will do: Spearhead…
…6+ years of experience in ML with a proven record of shipping large-scale models to production. Expertise in training and inference optimization. Proven…
…LLMs, deployment and distributed inference of LLMs, RAG, FM evaluation, Vector DBs, Agentic workflows, prompt/context engineering, and MLOps. - 6+ years of design/implementation…
…data, and increasingly AI/ML workloads. Many of these customers run accelerated compute at scale for distributed training and inference, so experience with GPU…
…LLMs, deployment and distributed inference of LLMs, RAG, FM evaluation, Vector DBs, Agentic workflows, prompt/context engineering, and MLOps. - Hands-on experience with AWS…
…optimization, dormancy detection, reactivation modeling). Define success metrics, experimentation frameworks (A/B, causal inference), and measurement methodology for seller engagement interventions. Productionize ML models…
…vision, large language models (LLMs), generative AI, causal inference, experimentation and A/B testing, optimization, and more! As an Applied Scientist Intern, you'll…
…with compiler engineers and runtime engineers to create, build and tune distributed inference solutions with Trainium and Inferentia. Experience optimizing inference performance for both…
…Hands-on experience with Python, PyTorch and GPU-accelerated training, inference, performance optimization, or serving. Research and engineering judgment, with the ability to find…
…Build and optimize the continuous autoregressive serving loop, including per-step control inputs, model and KV-cache state management, GPU inference, frame streaming, and…
…reasoning Spatial Multimodal Models Modality Alignment Model Scaling Synthetic data Inference Efficiency Inference Optimizations (e.g., parallel decoding, speculative decoding) Token Representations Model Distillation…
…training and inference Systems/AI algorithms codesign (e.g., for sparsity) AI/ML for systems (hardware design, code optimization, etc.) ML for EDA Programming…
…training and inference Systems/AI algorithms codesign (e.g., for sparsity) \ AI/ML for systems (hardware design, code optimization, etc.) ML for EDA Programming…
…Specialized expertise in Causal Inference at scale, Marketplace Optimization, or ML System Design for high-throughput environments. For New York City, NY-based roles…
…We're looking for ML engineers who want to define the product. That means less time on models in isolation and more on the…