Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer - GPU Fabric Observability”. A match may be a passing mention rather than the job itself. Titles only.
55 roles · group by role · page 1 of 3
…We are hiring a Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard…
…As a Quantum Systems Digital Design Engineer II , you’ll collaborate closely with quantum physicists, systems software, RF/EE, and manufacturing/production to define…
…As a Senior Supercomputing Operations Engineer, you will own day‑to‑day operations of GPU interconnect fabrics and treating them as a single, mission…
…fabrics are a first order reliability system that directly determines GPU availability, training throughput, and customer SLAs. As a Principal Supercomputing Operations Engineer, you…
…Engineering: Build and operate the high-performance networking fabrics, protocols, and observability needed for the largest training and serving workloads. - Hardware Health and Observability…
…About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in…
…Fabric & GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters…
…Engineering Culture NRB embraces automation and modern engineering practices. Engineers are encouraged to leverage AI-assisted development tools to accelerate software development, scripting, troubleshooting…
…Engineering Culture NRB embraces automation and modern engineering practices. Engineers are encouraged to leverage AI-assisted development tools to accelerate software development, scripting, troubleshooting…
…Lead operational and software engineering efforts that improve the reliability, availability, observability, and performance of OCI AI/HPC networking fabrics. Apply deep networking knowledge…
…As a Principal engineer, you will influence architectural direction, guide technical strategy, and collaborate across silicon, firmware, hardware, software, validation, manufacturing, and Azure engineering…
…stack. - Strong software engineering skills in Go (required) and Python; you write production-quality code, not just scripts - Deep experience with GPU orchestration in…
…as an SRE or DevOps engineer working with Kubernetes Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high…
…You’ll engineer ultra-low-latency ROCE networks, design congestion-free transport mechanisms, optimize lossless fabrics at 10k–100k+ GPU scale, and partner deeply…
…software - Engineer highly available network services through observability, failover, and redundancy features - Deliver predictable networking performance for clients via software-driven network engineering solutions…
…Observability and Telemetry teams to improve visibility into GPU health, network fabric performance, RDMA, RoCE, InfiniBand, and overall service health. Collaborate with internal engineering…
…across Crusoe Cloud’s GPU- and CPU-based infrastructure. You will be collaborating with hardware, software, infrastructure, and vendor engineering teams while working across…
…orchestration, and observability. Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack. Diagnosing…
…Unified observability plane. One pane that correlates GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals at extreme cardinality — so any engineer can diagnose…
…An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what…
…THE ROLE As a Software Engineer, Network Automation - Data Center Fabrics, you will help design and build systems for operating large-scale networks within…
…It encompasses various areas, including software and systems engineering practices, storage, data management, and services. Professionals in the role of Production Engineers hold specialized…
…across Crusoe Cloud’s GPU- and CPU-based infrastructure. You will be collaborating with hardware, software, infrastructure, and vendor engineering teams while working across…
…About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage…
…engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for…