Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Inference Engineer, GPU Kernel Optimization”. A match may be a passing mention rather than the job itself. Titles only.
105 roles across 118 listings · show every listing · page 4 of 5
…Develop detailed performance plans based on profiling findings and collaborate with NVIDIA's kernel engineering and OSS vLLM teams to drive improvements that benefit…
…Experience with CUDA kernel optimization and profiling. Experience with large-scale model training and production inference software stack. Strong collaborative and interpersonal skills, specifically…
…We are looking for a Senior Manager to lead the design, scaling, and operations of high-performance networking for GPU-based cloud infrastructure. This…
…across the full inference stack — from GPU kernels to framework-level bottlenecks — and ship measurable improvements. Work with the platform engineering side of the…
…software engineer to join our AI networking acceleration team, to work on a groundbreaking open-source library, using hardware offloads, GPU Kernels and RDMA…
…communication libraries for SoCs, ASICs, GPUs, CPUs, or FPGAs - Care about performance and have experience profiling and optimizing latency-sensitive or throughput-critical code…
…optimizing non-trivial GPU kernels. M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field. Strong understanding of GPU…
…kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) - Develop a suite of…
…for efficient inference Designing, implementing, and optimizing kernels for high impact AI workloads Designing and implementing extensible abstractions for LLM serving engines Building efficient…
…for efficient inference Designing, implementing, and optimizing kernels for high impact AI workloads Designing and implementing extensible abstractions for LLM serving engines Building efficient…
…are hiring a Senior Systems Software Engineer to join our team as a technical expert focused on optimizing deep learning inference for autonomous vehicles…
…GPU kernel authoring and performance analysis using tools such as Nsight Compute. A track record of success in mentoring early-career engineers and interns…
…software engineer to join our AI networking acceleration team, to work on a groundbreaking open-source library, using hardware offloads, GPU Kernels and RDMA…
…Mentor senior engineers and researchers, set technical direction, and raise the overall bar for systems rigor, performance engineering, and co-design thinking across the…
…on NVIDIA GPUs? We are now welcoming exceptional software engineers to apply to Senior Engineering positions in the Deep Learning Inference TensorRT software team…
…Our inference engine is designed to help developers run neural network models trained in a variety of frameworks on Snapdragon platforms at blazing speeds…
…Contribute to CUDA kernel and operator development for critical transformer components such as attention, GEMM, and MoE. Benchmark, profile, and optimize inference performance across…
…Familiarity with state-of-art neural network architectures, optimizers and LLM training. Experience with modern DL training frameworks and/or inference engines. Fluency in…
…Contributions to inference frameworks (vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe). Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies. Leading…
…Master's degree in Computer Science, Engineering, Information Systems, or related field 3+ years Hardware Engineering experience defining architecture of GPUs or accelerators used…
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization to…
We are now looking for a Senior Kernel Performance Architect for Deep Learning Software! NVIDIA is seeking extraordinary architects to develop processor and system…
…GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal…
…precision inference, quantization, compression of DNNs Experience with GPU programming Experience with building DSLs or optimizing compilers (e.g. graph compiler or kernel generator…
…inference systems. Familiarity with quantization, graph optimization, kernel fusion, and model partitioning. Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. Strong GPU…