Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Inference Engineer, GPU Kernel Optimization”. A match may be a passing mention rather than the job itself. Titles only.
105 roles across 118 listings · show every listing · page 2 of 5
The Local AI Automation team is seeking a Senior System Software Engineer passionate about building and maintaining robust system-level software and infrastructure for…
…That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting…
…performance inference Strong Python or C++ programming, software design, and software engineering skills. Hands-on experience with GPU kernel development or optimization (CUDA/C…
…This was for a senior engineer role. Please modify experience for engineer role Minimum Qualifications: • Bachelor's degree in Engineering, Information Systems, Computer Science…
…Preferred Skills - Experience with AI infrastructure, inference-serving systems, or large-scale machine learning systems. - Experience with compilers, runtimes, kernel optimization, or performance engineering…
…The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and…
…and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs. AI training and inference relies on…
…Senior Engineer for CoreWeave's Benchmarking & Performance team, focused on kernel authoring and optimization. You will write, profile, and tune the GPU kernels that…
…Experience with ML compilers and their internals, experience writing compiler optimization passes. Experience with accelerator HW architectures (TPUs/GPUs). Experience with ML Inference frameworks…
…GPU kernel authoring and performance analysis using tools such as Nsight Compute. A track record of success in mentoring early-career engineers and interns…
…Whether you're learning about Linux kernel internals, GPU driver architecture, Kubernetes device plugins, or distributed systems operations, our senior members provide one-on…
…Experience with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated components to handle high-bandwidth raw…
…Experience with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated components. Your base salary will be…
…storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams…
…This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads…
…Conduct in-depth GPU workload bottleneck analysis, implement system-level, kernel-level and framework-level tuning for AI training, inference, RL and gaming workloads…
…We are hiring dedicated Senior Software Engineers to join our GPU Fabric Networking group. This group builds and sustains system software that promotes swift…
…In addition, you will optimize ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on…
…We are now looking for an extraordinary Senior Perception Engineer to develop and productize NVIDIA’s autonomous driving solutions. As a member of our…
…Experience with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated components. Your base salary will be…
…Linux kernel and library optimizations for application performance tuning, managed Kubernetes implementations at scale, etc). Familiarity with advanced computing, AI, and/or GPU acceleration…
…Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving. - Maximize GPU Parallelism: Engineer and optimize GPU strategies…
…As the Engineering Manager for this team, you will lead a group focused on model optimization, training efficiency, GPU enablement, load testing, model performance…
…Build and enhance C++ & python backend implementations for ASR, TTS, and S2S pipelines, leveraging CUDA for GPU acceleration Optimize Inference Performance: Improve streaming latency…
…Have hands-on experience optimising PyTorch training or inference, profiling workloads, and reasoning about GPU memory, compute, and throughput. Are comfortable in containerised environments…