Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Inference Engineer, GPU Kernel Optimization”. A match may be a passing mention rather than the job itself. Titles only.
119 roles · group by role · page 1 of 5
…Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance…
…As a Software Engineer II or Senior Software Engineer - Simulation Platform, you will be responsible for designing, implementing, and ensuring quality of AI chip…
…Hands-on experience with Python, PyTorch and GPU-accelerated training, inference, performance optimization, or serving. Research and engineering judgment, with the ability to find…
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role…
…The engineer will also drive excellence in inference framework optimization and engineering best practices. The engineer will work closely with model development engineers, performance…
…Device drivers and kernel integration AI runtimes and execution engines with emphasis on graphs spanning multiple NPU. Compiler technologies and graph optimization ML frameworks…
…About the role As a Senior Principal Machine Learning Engineer, you will be responsible for designing, developing, and optimizing machine learning models—with a…
…Direct experience benchmarking or monitoring large GPU fleets or multi-region clusters. Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies…
…We are hiring an experienced kernel and performance engineer to take on this role at a senior level. You will own the performance of…
…We are hiring an experienced kernel and performance engineer to take on this work at a senior level. You will own the performance of…
…and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs. AI training and inference relies on…
Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem…
…kernel/Guest RDMA drivers. Industry-leading acceleration technologies such as TPUDirect, TPUDirect Storage (TDS), and GPUDirect Storage (GDS). As a Senior Staff Software Engineer…
…Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize the…
…Hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model…
…Familiarity with CUDA, Triton, or low-level GPU kernel development for inference pipeline acceleration. With a competitive salary package and benefits, NVIDIA is widely…
…Work alongside Staff and Senior engineers to resolve complex race conditions in the I/O path and optimize kernel-level memory pinning for GPU…
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We…
…mechanisms for GPU-intensive ML workloads. - Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers. Translate bottlenecks into…
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an…
…Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming, kernel optimization, and workload profiling Experience profiling…
…Engineering experience with LLM inference performance: profiling, kernel-level or serving-level optimization, or building a serving stack! Open-source contributions or product leadership…
…We work closely with ML researchers and developers to optimize and scale out model training and inference. The team operates at the intersection of…
…NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep learning—from LLMs and generative AI to recommendation, vision…
…Do you recognize Nitro, Trainium, EFA, ENA-X, Inferentia, Graviton? Join us and help create the future. We are seeking Senior SDE's interested…