Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Staff Engineer, Distributed Storage and HPC & AI Infrastructure”. A match may be a passing mention rather than the job itself. Titles only.
12 roles across 15 listings · show every listing
…in storage engineering, managing distributed storage at multi-petabyte scale Proven track record deploying and operating high-performance storage for GPU/HPC clusters Deep…
…post-sale, and serving as a critical technical voice between our customers and engineering teams. Ideal candidates are passionate about AI infrastructure, fluent in…
…Here, you’ll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate…
…RESPONSIBILITIES - Manage and mentor the systems and network engineering teams, providing leadership, technical direction, and career development for a distributed staff; help build and…
…RESPONSIBILITIES - Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads - Manage and optimize Slurm-based HPC environments for distributed…
…and storage product teams to ensure integrated infrastructure performance for distributed AI workloads - Define pricing and packaging models that balance utilization, customer flexibility, and…
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the…
…image distribution, and cold-start performance. Lead architecture and design reviews across infrastructure teams. Leadership and Influence Mentor senior and staff engineers and help…
…You will work at the intersection of DevOps, software engineering, and high-performance computing (HPC), building systems that accelerate chip design, simulation, and AI…
…AI/ML Infrastructure Expertise - Troubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures. - Support complex AI workloads (training + inference) with performance tuning and…
…generation infrastructure tooling, enabling Etched ASIC, Software, and Platform engineers to iterate faster, build more reliably, and push the boundaries of AI performance. This…
…helping scale infrastructure that supports demanding AI and HPC workloads. You'll partner closely with Production Engineers, infrastructure teams, and platform engineers to improve…