Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Reliability Engineer, Supercomputing”. A match may be a passing mention rather than the job itself. Titles only.
184 roles across 196 listings · show every listing · page 1 of 8
…ABOUT THE ROLE We're hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and…
…ready hardware for OpenAI’s supercomputing platform. ABOUT THE ROLE We’re looking for a Product Manufacturing Engineer to drive manufacturing strategy and execution…
…As a Senior Site Reliability Engineer, you will automate, and maintain large-scale distributed systems powering latest AI applications and machine learning models. Your…
…Engineering. • Simpler, better-controlled processes with measurable gains in efficiency, data quality, automation, and user experience. • Strong system adoption, sustained business ownership, and reliable…
…of-the-art AI models. • Collaborate with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale…
…As the engine that enables Microsoft’s cloud-first mission, CO+I delivers reliable, trusted, secure, and sustainable cloud and AI infrastructure to meet…
…At this supercomputing scale, reliability and operational excellence are engineering challenges of their own. As a Senior Supercomputing Operations Engineer, you will own day…
…a first order reliability system that directly determines GPU availability, training throughput, and customer SLAs. As a Principal Supercomputing Operations Engineer, you serve as…
…reliably. We are seeking a Developer Experience Engineer II and/or Senior Developer Experience Engineer - AI Frameworks, to enhance and accelerate our software engineering…
…driven engineers to establish both human and system safety controls throughout the expansion of SpaceX's large-scale AI infrastructure and supercomputing facilities in…
…This role works closely with engineering teams, data center technicians, vendors, and operations teams to bring new sites and infrastructure online quickly and reliably…
…This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability…
…Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and…
…generation of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are looking for an experienced software engineer to drive development for…
…reliability engineering. Ability to work collaboratively with cross-functional teams, including operations technicians. PREFERRED SKILLS AND EXPERIENCE: Experience in AI/ML infrastructure or supercomputing…
…reliability and efficiency targets Lead technical design reviews for network and system architecture changes affecting AI workload performance, communicating trade-offs clearly to engineering…
…THE ROLE We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute…
…reliable, high-impact research platform. You will lead all aspects of the day-to-day technical ownership of a tightly coupled GPU supercomputing environment…
…Are you a creative and autonomous engineer who loves a challenge? Are you ready to become the engineer you always wanted to be? Come…
…model capability. - Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs. - Build tools for yourself…
…As the engine that enables Microsoft’s cloud-first mission, CO+I delivers reliable, trusted, secure, and sustainable cloud and AI infrastructure, ensuring capacity…
…real-time, highly reliable system parts of leading Autonomous Vehicles. We are hiring now for the position of Senior Security Engineer, RTOS and Virtualization…
NVIDIA is looking for an experienced HPC DevOps Engineer to help us build the supercomputers and HPC clusters of the future. As a Senior…
…reliability. Successfully implementing and managing projects to meet ambitious deadlines and performance targets. What we need to see: BSc or MSc in Electrical Engineering…
…that enable fast initialization and reliable fault recovery for Ray and Kubernetes workloads running on next-generation GPU supercomputers. Our team develops the systems…