Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer, GPU Infrastructure (HPC)”. A match may be a passing mention rather than the job itself. Titles only.
86 roles across 101 listings · show every listing · page 1 of 4
…We are in search of a Deep Learning Software Infrastructure Engineer to propel NVIDIA’s Autonomous Vehicles project forward. In this role, you will…
…Our teams design and develop the hardware and software that connect infrastructure, accelerate computational workloads, integrate across the stack, protect data continuity, and deliver…
…These fabrics are the foundation underneath OCI's AI, GPU and HPC services, and support major tier-0 vendors in the generative AI industry…
…owns the software stack around NCCL (NVIDIA Collective Communications Library), which enables multi-GPU and multi-node data communication through HPC-style collectives. NCCL…
…minimizing instruction count and occupancy - Familiarity with HPC or ML infrastructure concepts - Experience with the full software development lifecycle including testing and operations Amazon…
…Our teams design and develop the hardware and software that connect infrastructure, accelerate computational workloads, integrate across the stack, protect data continuity, and deliver…
…Our teams design and develop the hardware and software that connect infrastructure, accelerate computational workloads, integrate across the stack, protect data continuity, and deliver…
…infrastructure, GPU clusters, HPC environments, or cloud networking platforms. Experience leading programs involving infrastructure automation, CI/CD, Infrastructure as Code, or cloud platform engineering…
…You will work closely with engineering to turn complex and ambiguous infrastructure problems into clear, durable product designs. You will define detailed behavior, examine…
…engineers build against. - Hands-on fluency with datacenter and GPU/HPC infrastructure — bare-metal provisioning (PXE/Redfish/IPMI), firmware and BIOS management, GPU health…
…OS-reprovisioning, or system-lifecycle tooling. - Have worked with GPU systems, AI infrastructure, HPC clusters, accelerators, or other heterogeneous and performance-sensitive compute environments…
…Experience within a world-class reliability function like Google SRE or Meta production engineering. Expertise in operating GPU, HPC, or AI training infrastructure with…
…GPU-accelerated solvers, HPC, or quantum annealing and QUBO formulations. Experience in hyper-growth startup-like environments, with demonstrated success balancing speed, ambiguity, and…
…RESPONSIBILITIES AI CLUSTER DEPLOYMENT & RACK INTEGRATION - Lead deployment of AI clusters including Cerebras Wafer Scale Engine, high-speed switches, storage, and rack-level infrastructure…
…Nice to Have Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards. Familiarity with AI training and…
…OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 3+ years Software Engineering, Hardware Engineering, Systems Engineering, or related…
…Crusoe Cloud is revolutionizing high-performance computing by offering sustainable, low-cost GPU compute power. As a Cloud Support Engineer, you'll play a…
…Crusoe Cloud is revolutionizing high-performance computing by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you'll play…
…We are looking for an Embedded Software Development engineer to build and own the server related firmware. As an embedded software development engineer in…
…The engineers who build the automation infrastructure for these packages don't just write software — they directly determine whether NVIDIA can tape out on…
…12+ overall years of experience in large scale storage architecture, operations, production engineering, or infrastructure. 6+ years of people management or technical leadership experience…
…Collaborate with scientists and software/infrastructure engineers to understand infrastructure requirements for training, testing, and deploying machine learning models. Implement automation solutions for provisioning…
…a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers. Drive the…
…AI infrastructure, unleashing the full compute bandwidth of clustered GPUs. About the Role: We are seeking a seasoned Staff Storage Software Engineer with deep…
…Specify node configurations and InfiniBand fabric options for clusters from 64 to 1,024+ GPUs, working with infrastructure engineering to keep NCCL (NVIDIA Collective…