Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Software Engineer, Deep Learning Inference - TensorRT”. A match may be a passing mention rather than the job itself. Titles only.
48 roles across 50 listings · show every listing · page 1 of 2
…are now welcoming exceptional software engineers to apply to Senior Engineering positions in the Deep Learning Inference TensorRT software team. What you’ll be…
…Machine Learning Engineer focused on research enablement and performance, turning promising experiments into stable, scalable, user-facing capabilities while making training and inference faster…
…You will collaborate with world-class engineers across deep learning software, compilers, GPU architecture, and open-source inference ecosystems, and your work will directly…
…Familiarity with NVIDIA software and deployment tools such as TensorRT, CUDA, cuDNN, Triton, DeepStream, TAO Toolkit, or RAPIDS. Experience building end-to-end pipelines…
…A love of research publications in the machine learning and software engineering communities Effective communicator with experience collaborating cross-functionally with other teams Additional…
…12+ years of software engineering experience in systems software, AI/ML infrastructure, deep learning inference, compiler/runtime technology, or platform performance. Strong C/C…
…Proficiency in Python and modern C++, with strong software engineering fundamentals (version control, testing, CI/CD). Deep understanding of 3D geometry, camera models, and…
We are now seeking a Senior Deep Learning Performance Architect! NVIDIA is looking for outstanding Performance Architects with a background in performance analysis, performance…
…We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads…
…Science, Machine Learning, or a related technical field preferred; Bachelor's considered with exceptional experience. 8+ years of professional software engineering experience in ML…
…12+ years in systems software engineering with hands-on experience in AI/ML workload optimization, GPU performance analysis, or deep learning infrastructure. Strong proficiency…
…Science, Electrical Engineering, or related field (or equivalent experience) and 12+ yrs of confirmed experience in systems software engineering with deep expertise in Windows…
…Engineering, Machine Learning, Robotics, or related field. 12+ years of experience in AI/ML systems, deep learning architecture, or hardware/software co-design. Deep…
…CRWV) in March 2025. Learn more at www.coreweave.com . About this role We’re looking for a Senior Engineer for CoreWeave’s Benchmarking…
We are seeking a Deep Learning Research Engineer to join our team and help develop the next generation of Large Language Model (LLM) inference…
…Experience with NVIDIA GPUs and software libraries, such as NVIDIA NeMo Framework , NVIDIA Triton Inference Server , TensorRT , TensorRT-LLM Excellent C/C++ programming skills…
…Excellent programming skills, particularly in Python and deep learning frameworks like PyTorch, and experience with software engineering standards. A strong problem-solving mentality and…
…We Prefer PhD in CS, EE, Deep Learning or a related field. Experience modifying ML compilers, runtimes, or inference engines (e.g., TensorRT, ONNX…
…Deep expertise in the performance internals and execution graphs of major deep learning autograd, training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM…
…Deep expertise in the performance internals and execution graphs of major deep learning autograd, training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM…
NVIDIA is hiring exceptional software engineers to build and optimize the core inference infrastructure for large language models. Join the TensorRT‑LLM team - the…
…deployment (quantization, pruning, NPU mapping) - Experience with AI accelerators and inference engines (TensorRT, ONNX Runtime, ExecuTorch, etc.) - Knowledge of camera control algorithms (AE, AWB…
…This role requires deep, hands-on fluency with open-source inference stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton), and…
…Deep learning familiarity: Experience with modern inference frameworks and an understanding of the architectural nuances of LLMs, Diffusion, and multi-modal models. Systems thinking…
…Write production-level, low latency, and memory-safe C++ and CUDA code for real-time inference on vehicle systems. Qualifications: Deep expertise in model…