Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Inference Engineer, GPU Kernel Optimization”. A match may be a passing mention rather than the job itself. Titles only.
105 roles across 118 listings · show every listing · page 3 of 5
…GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering…
…and inference system trends, especially around LLMs and generative AI. For senior candidates, strong experience in GPU kernel development and performance optimization, especially using…
…Advise labs on GPU-accelerated training, inference studies, agent evaluation, tool-use methods, data pipelines, scaling experiments, and reproducible workflows. Help build research prototypes…
…Build and scale our GPU infrastructure to support training and inference for frontier models, enabling innovation at scale while optimizing cost-to-serve. Stay…
…Familiarity with GPU/accelerator architecture and the complex scheduling challenges that come with it. Background in AI model development, training, inference A builder mindset…
…We are seeking a Senior Software Engineer for next-generation innovations in automotive platform performance, AI model optimization, scalability, and system architecture! In this…
…Knowledge and experience with Linux Operating System, administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging. Knowledge…
…Comfortable reading and modifying GPU kernels . Ways to stand out from the crowd: Experience optimizing LLM training or inference on multi-GPU NVIDIA systems…
…Coordinate GPU driver, display driver, and compute driver bring-up and validation on Windows (WDDM, MCDM) and Linux (open-gpu-kernel-modules, DRM/KMS…
…Benchmark the performance of state-of-the-art deep learning models’ inference and training passes to identify key GPU kernel and fusion opportunities. Identify…
…analyze, and optimize GPU compute kernels — targeting speed-of-light performance on NVIDIA hardware. Collaborate with GPU architects and performance engineers to encode domain…
…and CPU-to-GPU migration of scientific workloads. Perform low-level CUDA optimization, including custom kernels to accelerate simulation and inference workloads in drug…
…THE OPPORTUNITY - Own and operate GPU and accelerator clusters used for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration…
…MLIR to optimize high-level kernel descriptions (written in Triton's Python DSL), with a focus on generating efficient, low-level GPU code. When…
…Or specialized experience in runtime optimizations, model quantization, compression, on-device inference, GPU inference, pytorch, kernel development Your Location: This position is US - Remote…
…engines (e.g., TensorRT, ONNX Runtime, OpenXLA/PjRT, TVM). Experience building or scaling LLM serving systems, including expertise in distributed inference and performance optimization…
…You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi…
…versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference — from novel…
…GPU kernel generation with high performance and fast build time. A track record of success in mentoring junior engineers and interns is a bonus…
…What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipelines…
NVIDIA is hiring exceptional software engineers to build and optimize the core inference infrastructure for large language models. Join the TensorRT‑LLM team - the…
…This role requires deep, hands-on fluency with open-source inference stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton), and…
…Extreme performance optimization: Work at the intersection of Python orchestration and C++ engine-level optimizations to achieve major latency and throughput gains for critical…
…You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge…
We are looking for a software engineer with a strong background in parallel processing and GPU architecture to push the limits of performance at…