Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Reliability Engineer, Supercomputing”. A match may be a passing mention rather than the job itself. Titles only.
27 roles across 29 listings · show every listing · page 1 of 2
…ABOUT THE ROLE We're hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and…
…As a Software Engineer on the Cluster Deployment Automation team, you will help build the pushbutton tooling that makes large-scale cluster deployments faster…
NVIDIA is a leading artificial intelligence computing company, and we are paving the way with innovations in self-driving cars, machine learning, supercomputing, gaming…
…Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity…
…providers to enable rapid, reliable deployment of NVIDIA solutions globally. Mentor and guide Infra SA team members and partner engineering teams, sharing best practices…
…hard problems like supercomputing, high availability, and building loosely coupled gRPC-based microservices all while continuing to securely and reliably handling massive traffic with…
…scale across our supercomputing infrastructure, both on prem and in the cloud. You will work cross-functionally across ML engineering, data science, and research…
…As senior software development engineer in Test, we are looking for a candidate who can make a big impact on how we test and…
…engineers daily using extensive knowledge of network topologies, physical and logical, and network protocols. Expert knowledge and proven history with designing scalable and reliable…
…As senior software development engineer in Test, we are looking for a candidate who can make a big impact on how we test and…
…telemetry, and log analysis - Exposure to production systems validation or infrastructure reliability engineering Why Join Cerebras People who are serious about software make their…
…productivity, infrastructure, and engineering acceleration. - Mentor engineers, review designs, guide architecture, and set a high bar for engineering quality, reliability, and operational excellence. - Own…
…Minimum Skills & Qualifications - 3+ years of experience in software engineering, QA/quality engineering, systems engineering, or infrastructure development. - Strong programming skills in Python and…
…Responsibilities: - Contribute to the development and maintenance of CICD pipelines, ensuring reliable and efficient build, test, and release workflows across the organization. - Help manage…
…Security and Engineering teams to deliver secure-by-design solutions, implement Zero Trust principles, and reduce operational friction. - Write high-quality, reliable code, participate…
…commissioned, and handed over to meet Cerebras' performance, reliability, scalability, and schedule requirements. The Mechanical Engineer will serve as the owner's technical representative…
…Track record leading bring-up and deployment of large clusters or supercomputing environments. Background in external customer-facing roles (field engineering, escalations, or pre…
…providers and enterprises to deliver high-performance, reliable, and secure networking solutions. Working alongside exceptional engineers across the globe, you'll thrive in a…
…Deep understanding of congestion management, QoS, security isolation Deep understanding of networking reliability, availability and serviceability Capable of abstracting and analyzing network performance from…
…Required Experience & Skills - 15+ years in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at…
…largest AI supercomputers on Earth. RESPONSIBILITIES: Develop routing and traffic-engineering algorithms for the Colossus high-performance datacenter network. Develop highly reliable, real-time…
…BS or MS in Engineering, Electrical Engineering, Physics, or Computer Science (or equivalent experience). 5+ years of work-related experience in high-tech IT…
…that support SpaceXAI's AI supercomputing goals. You will serve as a key bridge between field execution, engineering, and executive leadership, driving program-level…
…performance, and reliability targets. - Own incident response coordination for facility-related issues. - Maintain alignment with global: - Cluster Operations; - Network Operations; - Platform Engineering. FACILITY INTERFACE…
…This role sits within the AI Infrastructure organization and is responsible for ensuring that high-density AI clusters are built rapidly, operated reliably, and…