Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Reliability Engineer, Supercomputing”. A match may be a passing mention rather than the job itself. Titles only.
184 roles across 196 listings · show every listing · page 2 of 8
…This is a hands-on software engineering role at the boundary of distributed systems and AI supercomputing. You will design control planes, APIs, workflow…
…Research and integrate sophisticated software engineering practices, automation tools, and generative AI technologies to improve software reliability, maintainability, and scalability. What we need to…
…engineer to design, build, and operate the GPU supercomputing environment that powers large‑scale training and inference. You will deliver high‑performant, reliable, and…
…They translate the high-level designs provided by the Customer Success organization into fully operational supercomputers. In conjunction with the Datacenter Engineering team, they…
…driven engineers to establish both human and system safety controls throughout the expansion of SpaceX's large-scale AI infrastructure and supercomputing facilities in…
…Build and maintain the lean, high-reliability Linux-based operating system that underpins our supercomputer network fabric. Write (or rewrite) high-performance device drivers…
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another…
NVIDIA is a world‑leading, fast‑growing AI computing company, delivering everything from the most powerful GPU‑accelerated supercomputers to gigawatt‑scale AI data…
…the infrastructure powering NVIDIA AI supercomputers. In this role, you will manage and develop a team of talented engineers while remaining closely connected to…
…clusters or supercomputing environments. Background working with fast-moving AI labs or frontier-model infrastructure deployments and customer-facing roles (field engineering, or pre…
…BS or MS degree in EE/CS/CE (or equivalent experience) 12+ years of infrastructure, platform, DevOps, or SRE engineering Deep expertise in modern…
…clusters or supercomputing environments. Background working with fast-moving AI labs or frontier-model infrastructure deployments and customer-facing roles (field engineering, or pre…
…Responsibilities Advanced ROCE transport design, congestion control, ECN/WRED/DCTCP tuning Fabric architecture, topology planning, network modeling, and scaling strategy Telemetry, observability, reliability engineering…
…safe, reliable, scalable, and repeatable data center designs, working closely with internal manufacturing partners and Cerebras mechanical, thermal, hardware, controls, and systems engineering teams…
…SKILLS & QUALIFICATIONS MINIMUM - Bachelor's or master’s degree in mechanical engineering, Electrical Engineering, Architectural Engineering or related field. - 10+ years of mission-critical…
…We are looking for a versatile Electrical Engineer who thrives at the bench—diagnosing complex, multi-domain failures and driving root cause to the…
…Turn one-off investigations into repeatable engineering systems. MINIMUM QUALIFICATIONS - 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production…
…Hardware Analytics Engineer Job Duties: - Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization. - Architect, develop…
…Engineer Job Duties: - Design and develop automated test frameworks and execute software validation using Python, Java, and Selenium to ensure quality and reliability of…
…The Production Engine for Inference Core — helping turn integrated features into reliable production releases. You will write software and automation, test new model and…
…The Production Engine for Inference Core — turning integrated features into reliable production releases. You will define the quality strategy across the pre-release and…
…SKILLS & QUALIFICATIONS - 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering. - Hands-on experience building or…
…execution engines, scheduling systems, test infrastructure, developer tools, and reusable software platforms that allow engineers to build, test, qualify, and deliver software reliably at…
…As a Software Engineer on the Cluster Deployment Automation team, you will help build the pushbutton tooling that makes large-scale cluster deployments faster…
NVIDIA is a leading artificial intelligence computing company, and we are paving the way with innovations in self-driving cars, machine learning, supercomputing, gaming…