Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Reliability Engineer, Supercomputing”. A match may be a passing mention rather than the job itself. Titles only.
38 roles · page 1 of 2
…This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability…
…Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and…
…generation of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are looking for an experienced software engineer to drive development for…
…reliability engineering. Ability to work collaboratively with cross-functional teams, including operations technicians. PREFERRED SKILLS AND EXPERIENCE: Experience in AI/ML infrastructure or supercomputing…
…reliability and efficiency targets Lead technical design reviews for network and system architecture changes affecting AI workload performance, communicating trade-offs clearly to engineering…
…THE ROLE We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute…
…reliable, high-impact research platform. You will lead all aspects of the day-to-day technical ownership of a tightly coupled GPU supercomputing environment…
…Are you a creative and autonomous engineer who loves a challenge? Are you ready to become the engineer you always wanted to be? Come…
…model capability. - Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs. - Build tools for yourself…
…As the engine that enables Microsoft’s cloud-first mission, CO+I delivers reliable, trusted, secure, and sustainable cloud and AI infrastructure, ensuring capacity…
…real-time, highly reliable system parts of leading Autonomous Vehicles. We are hiring now for the position of Senior Security Engineer, RTOS and Virtualization…
NVIDIA is looking for an experienced HPC DevOps Engineer to help us build the supercomputers and HPC clusters of the future. As a Senior…
…reliability. Successfully implementing and managing projects to meet ambitious deadlines and performance targets. What we need to see: BSc or MSc in Electrical Engineering…
…that enable fast initialization and reliable fault recovery for Ray and Kubernetes workloads running on next-generation GPU supercomputers. Our team develops the systems…
…This is a hands-on software engineering role at the boundary of distributed systems and AI supercomputing. You will design control planes, APIs, workflow…
…Research and integrate sophisticated software engineering practices, automation tools, and generative AI technologies to improve software reliability, maintainability, and scalability. What we need to…
…engineer to design, build, and operate the GPU supercomputing environment that powers large‑scale training and inference. You will deliver high‑performant, reliable, and…
…They translate the high-level designs provided by the Customer Success organization into fully operational supercomputers. In conjunction with the Datacenter Engineering team, they…
…driven engineers to establish both human and system safety controls throughout the expansion of SpaceX's large-scale AI infrastructure and supercomputing facilities in…
…Build and maintain the lean, high-reliability Linux-based operating system that underpins our supercomputer network fabric. Write (or rewrite) high-performance device drivers…
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another…
NVIDIA is a world‑leading, fast‑growing AI computing company, delivering everything from the most powerful GPU‑accelerated supercomputers to gigawatt‑scale AI data…
…the infrastructure powering NVIDIA AI supercomputers. In this role, you will manage and develop a team of talented engineers while remaining closely connected to…
…clusters or supercomputing environments. Background working with fast-moving AI labs or frontier-model infrastructure deployments and customer-facing roles (field engineering, or pre…
…BS or MS degree in EE/CS/CE (or equivalent experience) 12+ years of infrastructure, platform, DevOps, or SRE engineering Deep expertise in modern…