Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Inference Engineer, GPU Kernel Optimization”. A match may be a passing mention rather than the job itself. Titles only.
60 roles across 67 listings · show every listing · page 1 of 3
…As the Engineering Manager for this team, you will lead a group focused on model optimization, training efficiency, GPU enablement, load testing, model performance…
…Build and enhance C++ & python backend implementations for ASR, TTS, and S2S pipelines, leveraging CUDA for GPU acceleration Optimize Inference Performance: Improve streaming latency…
…Have hands-on experience optimising PyTorch training or inference, profiling workloads, and reasoning about GPU memory, compute, and throughput. Are comfortable in containerised environments…
…GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering…
…and inference system trends, especially around LLMs and generative AI. For senior candidates, strong experience in GPU kernel development and performance optimization, especially using…
…Advise labs on GPU-accelerated training, inference studies, agent evaluation, tool-use methods, data pipelines, scaling experiments, and reproducible workflows. Help build research prototypes…
…Build and scale our GPU infrastructure to support training and inference for frontier models, enabling innovation at scale while optimizing cost-to-serve. Stay…
…Familiarity with GPU/accelerator architecture and the complex scheduling challenges that come with it. Background in AI model development, training, inference A builder mindset…
…We are seeking a Senior Software Engineer for next-generation innovations in automotive platform performance, AI model optimization, scalability, and system architecture! In this…
…Knowledge and experience with Linux Operating System, administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging. Knowledge…
…Comfortable reading and modifying GPU kernels . Ways to stand out from the crowd: Experience optimizing LLM training or inference on multi-GPU NVIDIA systems…
…Coordinate GPU driver, display driver, and compute driver bring-up and validation on Windows (WDDM, MCDM) and Linux (open-gpu-kernel-modules, DRM/KMS…
…Benchmark the performance of state-of-the-art deep learning models’ inference and training passes to identify key GPU kernel and fusion opportunities. Identify…
…analyze, and optimize GPU compute kernels — targeting speed-of-light performance on NVIDIA hardware. Collaborate with GPU architects and performance engineers to encode domain…
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We…
…and CPU-to-GPU migration of scientific workloads. Perform low-level CUDA optimization, including custom kernels to accelerate simulation and inference workloads in drug…
…Experience with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated components.
…THE OPPORTUNITY - Own and operate GPU and accelerator clusters used for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration…
…MLIR to optimize high-level kernel descriptions (written in Triton's Python DSL), with a focus on generating efficient, low-level GPU code. When…
…Or specialized experience in runtime optimizations, model quantization, compression, on-device inference, GPU inference, pytorch, kernel development Your Location: This position is US - Remote…
…engines (e.g., TensorRT, ONNX Runtime, OpenXLA/PjRT, TVM). Experience building or scaling LLM serving systems, including expertise in distributed inference and performance optimization…
…You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi…
…Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming and kernel optimization. A strong analytical approach…
…versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference — from novel…
…GPU kernel generation with high performance and fast build time. A track record of success in mentoring junior engineers and interns is a bonus…