Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Staff Engineer, Distributed Storage and HPC & AI Infrastructure”. A match may be a passing mention rather than the job itself. Titles only.
26 roles · group by role · page 1 of 2
…in storage engineering, managing distributed storage at multi-petabyte scale Proven track record deploying and operating high-performance storage for GPU/HPC clusters Deep…
…Industry-leading acceleration technologies such as TPUDirect, TPUDirect Storage (TDS), and GPUDirect Storage (GDS). As a Senior Staff Software Engineer, you will lead the…
…Staff Software Engineer to build, scale, and operate our custom High-Performance Computing infrastructure. As Zoox scales its autonomous vehicle development, our HPC platform…
…will shape the infrastructure that powers the next generation of AI training and inference at scale. As a Staff Engineer on our Orchestration team…
…Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About…
…of Technical Staff, Software Engineers - Compute Infra /HPC to build the compute infra software that brings frontier AI compute online and keeps it healthy…
…distributed compute platform. - Experience supporting GPU, HPC, or large-scale AI training infrastructure. - Experience with distributed storage, cluster schedulers, cloud providers, or infrastructure control…
…and scalability in AI/ML and HPC workloads. AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In…
…and track storage SLOs/SLIs, and ensure reliable deployment and maintenance of distributed storage infrastructure. - Innovate: Stay current with AI and HPC storage research…
…running distributed AI/ML/HPC workloads across thousands of GPUs, leveraging technologies like RoCE and Infiniband. As a Consulting Member of Technical Staff, you…
…infrastructure for AI/ML or HPC workloads. - Hands-on experience with distributed training and/or inference workloads at scale, including parallelism strategies and performance…
…and implement high-speed networking infrastructure (100G/200G/400G Ethernet) connecting compute nodes, storage, and upstream peering, in coordination with network and platform engineering…
…post-sale, and serving as a critical technical voice between our customers and engineering teams. Ideal candidates are passionate about AI infrastructure, fluent in…
…Here, you’ll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate…
…Here, you’ll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate…
…Here, you’ll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate…
…RESPONSIBILITIES - Manage and mentor the systems and network engineering teams, providing leadership, technical direction, and career development for a distributed staff; help build and…
…RESPONSIBILITIES - Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads - Manage and optimize Slurm-based HPC environments for distributed…
…As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale…
…and storage product teams to ensure integrated infrastructure performance for distributed AI workloads - Define pricing and packaging models that balance utilization, customer flexibility, and…
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the…
…image distribution, and cold-start performance. Lead architecture and design reviews across infrastructure teams. Leadership and Influence Mentor senior and staff engineers and help…
…You will work at the intersection of DevOps, software engineering, and high-performance computing (HPC), building systems that accelerate chip design, simulation, and AI…
…AI/ML Infrastructure Expertise - Troubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures. - Support complex AI workloads (training + inference) with performance tuning and…
…generation infrastructure tooling, enabling Etched ASIC, Software, and Platform engineers to iterate faster, build more reliably, and push the boundaries of AI performance. This…