Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer, GPU Inference”. A match may be a passing mention rather than the job itself. Titles only.
170 roles across 179 listings · show every listing · page 1 of 7
…We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working across our custom inference APIs, the vLLM serving runtime…
…own inference services - Support and debug production issues through on-call rotation Required Qualifications - Have 6+ years of experience in software engineering, with a…
…complex software systems and continuously improve the quality of other AI agents. ABOUT THE ROLE: We’re looking for an experienced research engineer to…
…As a Software Engineer II on the team, you will help design, build, and harden a secure, evaluable, vendor-agnostic agent platform that teams…
…1) AI Infrastructure with purpose-built chips like AWS Trainium and Inferentia, GPU-powered instances, and optimized frameworks like PyTorch and JAX, 2) AI…
By submitting your resume, you acknowledge that your 2027 Software Engineering internship application will be processed in accordance with NVIDIA’s Applicant Privacy Policy…
…are optimized for Training and Inferencing. Lead the development of end-to-end AI network architecture including gpu-gpu, server – storage, in-rack cabling…
…Experience building or supporting production AI/ML platforms (training, deployment, and model serving/inference), including GPU infrastructure/tooling. Strong DevOps/platform engineering practices: CI…
…You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers…
…software engineering with deep specialization in compiler development, code generation, or performance optimization for accelerators Experience architecting production compiler infrastructure for ML accelerators, GPUs…
…GPU utilization and efficiency, interconnect (NVLink, InfiniBand, RoCE) fabric behavior, distributed training and inference throughput, and hardware degradation signals. Build the detection and diagnosis…
…Leadership & Strategy: - Lead, mentor, and grow a team of high-caliber software engineers on Crusoe’s - Partner with leadership to define and execute the…
…Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale AI training and inference clusters. In…
…generative AI inference and training as successive silicon generations land (see [https://bit.ly/metamtia](https://bit.ly/metamtia)). The MTIA Software team is…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…Collaborate with AI Infra, Hardware Engineering, Product, and Research; represent the team to leadership 3+ years managing software engineering teams with a track record…
…Experience with inference-serving frameworks, GPU-aware scheduling, or model-performance optimization. Expertise in vector databases, GPU-accelerated query engines, or distributed data platforms…
…for next-generation NVIDIA GPUs. Advance the state-of-the-art: Solve complex compilation problems for AI workloads (both inference and training) and successfully…
…Engineering Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads. Build, deploy, and operate components supporting LLM inference, agentic…
…and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs. AI training and inference relies on…
…Proven ability to design, implement, and optimize scalable ML architectures, from distributed training to real-time inference. Strong software engineering skills in Python, C…
…LLM inference, an inter-agent message bus, parametric CAD pipelines, a print farm, and the infrastructure connecting it all. We need an engineer who…
…LLM inference, an inter-agent message bus, parametric CAD pipelines, a print farm, and the infrastructure connecting it all. We need an engineer who…
…develop the compiler for Apple's proprietary Neural Engine Accelerator, optimizing it for deep learning inference with a focus on performance, scalability, and power…
…key workloads with ultra high-speed inference. Key Responsibilities - Work with architects, designers, post silicon and software engineers to ensure a high-quality design…