Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer - GPU Fabric Observability”. A match may be a passing mention rather than the job itself. Titles only.
13 roles across 14 listings · show every listing
…Engineering: Build and operate the high-performance networking fabrics, protocols, and observability needed for the largest training and serving workloads. - Hardware Health and Observability…
…About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in…
…Fabric & GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters…
…Engineering Culture NRB embraces automation and modern engineering practices. Engineers are encouraged to leverage AI-assisted development tools to accelerate software development, scripting, troubleshooting…
…Lead operational and software engineering efforts that improve the reliability, availability, observability, and performance of OCI AI/HPC networking fabrics. Apply deep networking knowledge…
…As a Principal engineer, you will influence architectural direction, guide technical strategy, and collaborate across silicon, firmware, hardware, software, validation, manufacturing, and Azure engineering…
…stack. - Strong software engineering skills in Go (required) and Python; you write production-quality code, not just scripts - Deep experience with GPU orchestration in…
…as an SRE or DevOps engineer working with Kubernetes Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high…
…You’ll engineer ultra-low-latency ROCE networks, design congestion-free transport mechanisms, optimize lossless fabrics at 10k–100k+ GPU scale, and partner deeply…
…software - Engineer highly available network services through observability, failover, and redundancy features - Deliver predictable networking performance for clients via software-driven network engineering solutions…
…Observability and Telemetry teams to improve visibility into GPU health, network fabric performance, RDMA, RoCE, InfiniBand, and overall service health. Collaborate with internal engineering…
…across Crusoe Cloud’s GPU- and CPU-based infrastructure. You will be collaborating with hardware, software, infrastructure, and vendor engineering teams while working across…
…orchestration, and observability. Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack. Diagnosing…