Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer, Inference & Optimization”. A match may be a passing mention rather than the job itself. Titles only.
271 roles across 296 listings · show every listing · page 1 of 11
…5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale. - Inference Mastery: Proven expertise in inference optimization…
…a formal engineering technical leadership role, leading software engineering teams. 2 years of experience in LLM training or inference, including performance optimizations, distributed execution…
…engineering management role, guiding software engineering teams, including hiring or team development. 2 years of experience in LLM training or inference, including performance optimizations…
…Expert-level software engineering fundamentals using Python, PyTorch, or JAX, with a track record of building reliable, highly scalable ML systems. Proven ability to…
…Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is…
…Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is…
…Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is…
…MINIMUM QUALIFICATIONS: - Bachelor’s degree in Computer Science, Engineering, or a related technical field. - 5+ years of experience in a software engineering role, with…
…with an ML engineer in the same afternoon. Preferred Qualifications - 10+ years in technical field or engineering roles. - Experience with inference serving frameworks (vLLM…
…MINIMUM QUALIFICATIONS: - Bachelor’s degree in Computer Science, Engineering, or a related technical field. - 5+ years of experience in a software engineering role, with…
…of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - Experience developing using AI tooling such as Claude Code…
…Building and optimizing local AI inference stack for RTX, RTX Pro and DGX GPUs, focusing on performance, stability, and scalability across various hardware architectures…
…optimization, deployment, and production inference. The ideal candidate is a senior technical DevRel leader who can earn credibility with ML researchers and platform engineers…
…Benchmarking & Performance Engineering Optimize CPU-centric ML benchmarks such as: Geekbench AI Internal benchmarking suites Establish performance baselines and track improvements across hardware generations…
…multimodal systems, and computation optimizations. –Designs and reviews the architecture of ML and AI solutions, including data, model, training, inference, and evaluation components, employing…
…with an ML engineer in the same afternoon. Preferred Qualifications - 10+ years in technical field or engineering roles. - Experience with inference serving frameworks (vLLM…
…and inference concepts clearly, including memory bandwidth, interconnect, quantization, throughput vs. latency, and cost-per-token tradeoffs Partner closely with product, engineering, developer relations…
…machine learning systems and/or platforms. - Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them…
…You’ll primarily partner with machine learning engineering leaders to deliver high-impact science and analytics that power decision automation and optimize performance across…
…The team is also responsible for the infrastructure for running ML inference across fleets of devices. We are building the first end-to-end…
…Crucially, you’ll take the friction you encounter in the field—like optimizing inference speeds, resolving interoperability issues, and connecting complex external data sources…
…Crucially, you’ll take the friction you encounter in the field—like optimizing inference speeds, resolving interoperability issues, and connecting complex external data sources…
…High speed networking (RDMA), Distributed ML Training, GPU architecture, ML systems, AI infrastructure, high performance computing, performance optimizations, or Machine Learning frameworks (e.g…
…Experience working with real-time ML inference, A/B testing, and optimization frameworks. Experience translating ML evaluation results and performance metrics into actionable product…
…of ML converters/compilers and runtimes, and hardware-accelerated ML inference techniques. Strong understanding of generative AI model architectures and their optimization for on…