Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Compiler Engineer - AI Inference”. A match may be a passing mention rather than the job itself. Titles only.
296 roles across 342 listings · show every listing · page 3 of 12
…operator fusion, scheduling, quantisation-aware performance, custom kernels) Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory…
Overview The Microsoft AI Frameworks team, develops the software , performance systems, and engineering tools that enable state-of-the-art AI models to run…
…You will work with applied scientists, ML engineers, GPU kernel engineers, compiler and runtime teams, hardware teams, and product teams to deliver systems for…
…bottlenecks in both training and inference pipelines. Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts…
…bottlenecks in both training and inference pipelines. Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts…
…visualization tools that help engineers identify performance bottlenecks and optimize models. - Work alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and…
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme…
…On this team, you’ll build an MLIR-based AI compiler that powers NVIDIA’s inference engine end to end, with a focus on…
…and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe…
…As an AI Engineer on our team, you will own the infrastructure and tooling that let LLM-powered features ship reliably at Apple scale…
…Job Summary Etched is building a new category of AI hardware: frontier inference clusters. As a Physical Design Engineer, you will own block-level…
…platforms - Proven track record in building AI agents that automate ML workload optimization, ML compiler tuning, distributed inference and training, or ML kernel authoring…
…pruning, knowledge distillation) and efficient inference algorithm Strong background on compiler stack and ML system optimization for AI accelerators (e.g., graph transformation, graph…
…software) for inference or training solutions. Develops optimized software to enable AI models deployed on hardware (e.g., machine learning kernels, compiler tools, or…
…the AI software stack, including the fundamental abstractions, programming models, compilers, runtimes, libraries and APIs to enable large scale training and inferencing of models…
…You will work closely with AI/ML Scientists and engineers at the intersection of Generative AI and Information Retrieval, crafting intelligent systems that personalize…
…Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both…
…This is a hands-on compiler engineering leadership role for someone who understands where compiler quality can regress in modern AI compiler stacks. What…
…Experience developing code with agents — Claude Code, Cursor, Open AI and inference SDKs. Background integrating compilers and developer tools with AI coding agents or…
…and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe…
We are now looking for a Software Development Engineer for LLM inference! NVIDIA is hiring software engineers for its TensorRT-LLM team. Academic and…
NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment…
…stack, including compiler, model onboarding, inference serving, platform software, ecosystem enablement, firmware, systems, and hardware. You will work closely with AI engineering teams to…
…Work across AI frameworks, runtime, compiler, kernels, distributed systems, infrastructure, and hardware. - Inference-path focus: Learn to validate high-risk changes across runtime, host…
…Release Integration Testing within Release & Feature Qualification for AI Inference Core. The Production Engine for Inference Core — turning integrated features into reliable production releases…