Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer, Inference & Optimization”. A match may be a passing mention rather than the job itself. Titles only.
1,291 roles across 1,496 listings · show every listing · page 2 of 52
…Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution - Experience using managed ML/AI solutions Amazon is an equal opportunity…
…or Storage - • Experience helping customers design and optimize infrastructure for AI/ML workloads (training clusters, inference optimization, GPU scheduling, etc.) - *Demonstrated ability to adapt…
…Partners with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns. Optimizes platform reliability…
…Experience with GPU inference optimization, model serving, workflow engines, or multi-tenant cloud services. Experience integrating AI systems with enterprise APIs, databases, identity systems…
…Our goal is to continuously improve our custom enterprise models by incorporating novel techniques and optimizing training and inference efficiency, establishing a differentiated advantage…
…You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take…
…We work closely with ML researchers and developers to optimize and scale out model training and inference. The team operates at the intersection of…
…7 years of experience leading technical project strategy, ML design, and optimizing industry ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging…
…compute clusters optimized for AI inference and training workloads at the tactical edge (NVIDIA H100/A100, AMD MI300, or similar) Engineer advanced thermal solutions…
…About this role As the Senior Staff Machine Learning Platform Engineer, you will own the technical vision and evolution of Faire’s ML platform…
…Raise the bar for engineering excellence: reliability, performance, security/privacy, cost efficiency, and operational health of large-scale ML systems. Represent Pinterest externally for…
…They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability…
…Apply AI/ML methods where they create practical leverage, including predictive modeling, optimization, and LLM-enabled workflows that improve decision speed, quality, or scale…
…with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale AI training and inference systems. • Build…
…Establish and evolve architectural standards, patterns, and best practices across the SambaStack engineering organization. Partner with ML Engineers, Product Managers, QA, and DevOps to…
…As a Staff Engineer, you will own the technical roadmap for our ML platform. You will build the robust infrastructure, MLOps tooling, and systems…
…ML services (e.g., Bedrock/SageMaker, Azure OpenAI, Vertex AI) Experience with Kubernetes and containerization is welcome; experience serving models or deploying inference/GPU…
…ML services (e.g., Bedrock/SageMaker, Azure OpenAI, Vertex AI) Experience with Kubernetes and containerization is welcome; experience serving models or deploying inference/GPU…
…of LLM architectures and modern inference engines like vLLM, TensorRT-LLM, or SGLang - The ability to profile and optimize GPU workloads, in training or…
…models for domain-specific tasks, and prompt engineering optimized for performance, reliability, and safety. Experience with MLOps and LLMOps: model lifecycle management, deployment pipelines…
…Architect production-ready inference solutions leveraging Mantle's distributed engine, Zero Operator Access security model, and OpenAI-compatible APIs—optimizing for customer-specific requirements…
…LLMs, deployment and distributed inference of LLMs, RAG, FM evaluation, Vector DBs, Agentic workflows, prompt/context engineering, and MLOps. - Hands-on experience with AWS…
…vision, large language models (LLMs), generative AI, causal inference, experimentation and A/B testing, optimization, and more! As an Applied Scientist Intern, you'll…
…About the team The Hardware Engineering AI/ML development team is a group of engineers and technical program managers directly responsible for launching and…
…ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT - Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles…