Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer, GPU Infrastructure (HPC)”. A match may be a passing mention rather than the job itself. Titles only.
389 roles · group by role · page 1 of 16
…As a Staff Software Engineer, you will: - Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple…
…About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime…
…our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The Fleet Engineering teams: - HPC Deployments — Turns…
…owns the software stack around NCCL (NVIDIA Collective Communications Library), which enables multi-GPU and multi-node data communication through HPC-style collectives. NCCL…
…About the team The team is comprise of both Hardware Design Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with…
…12+ years of professional software engineering experience, with 8+ years in service operations, monitoring, and reliability improvement for infrastructure. 5+ years of hands-on…
…Compute, Generative AI, Accelerated Compute/GPU infrastructure, HPC, Security, Networking, or Storage - • Experience helping customers design and optimize infrastructure for AI/ML workloads (training…
…PDU/UPS suppliers), or in hyperscale/colocation infrastructure engineering. Hands-on experience deploying liquid-cooled AI/HPC clusters or high-density rack solutions at…
…Infrastructure) AI Infrastructure is at the forefront of building a cutting-edge, ultra-high-performance GPU platform designed to support AI/ML/HPC workloads…
…management, infrastructure engineering, or related technical leadership roles. 1+ years experience building or managing large-scale cloud, HPC, GPU, AI, or datacenter infrastructure deployments…
…in Computer Engineering, Computer Science, Electrical Engineering, or related technical field 5+ years of experience designing and deploying data center infrastructure, HPC clusters, large…
…3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles. Experience operating distributed systems in production and keeping…
…level infrastructure issues Demonstrated ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve production issues #azurecorejobs Software Engineering IC4…
…HPC), or artificial intelligence (AI) infrastructure in production environments Demonstrated ownership of mission‑critical production infrastructure with direct impact on service availability, GPU workloads…
…HPC and networking. We are well positioned as the “AI Computing Company”, and our GPUs are the brains powering modern Deep Learning software frameworks…
…in Electrical Engineering, Computer Engineering, Computer Science, or equivalent practical experience 12+ years architecting hardware systems for hyperscale, HPC, or AI/ML infrastructure Deep…
…experience in a technical role (e.g., systems engineering, DevOps, ML infrastructure, solutions architecture, or software development) - Experience with operational parameters and troubleshooting for…
…preferably for high density AI/HPC data centers. Proven experience delivering large-scale data center hardware deployments or infrastructure planning programs, at multi-rack…
NVIDIA is the world leader in GPU Computing. We are passionate about four markets: Gaming, Automotive, Enterprise Graphics and HPC/Cloud Datacenters; in addition…
Design, develop, troubleshoot and debug software programs for GPU-based AI infrastructure OCI (Oracle Cloud Infrastructure) AI Infrastructure is at the forefront of building…
…projects within software development teams, with experience delivering infrastructure platforms Experience with high-performance GPU concepts such as RDMA, RoCE and HPC concepts more…
…Hands-on experience with HPC clusters, InfiniBand, GPU infrastructure, or hyperscale data center technologies. Experience in AI infrastructure deployment, professional services, or tech vendor…
NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers and…
…We are seeking motivated, personable, and independent individuals to join our team! We seek experienced software embedded engineers to help support our groundbreaking, innovative…
…Senior Software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes…