Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer - GPU Fabric Observability”. A match may be a passing mention rather than the job itself. Titles only.
52 roles across 55 listings · show every listing · page 2 of 3
…and corporate environments. - Set standards for resiliency, observability, and capacity planning across the global network - Mentor engineers across regions (US, Canada, Bangalore, and beyond…
…Familiarity with GPU and fabric telemetry (e.g., DCGM, NVLink, InfiniBand/Ethernet fabric counters) and using it to diagnose performance regressions. Strong communication skills…
…Lead operational and software engineering efforts that improve the reliability, availability, observability, and performance of OCI AI/HPC networking fabrics. Apply deep networking knowledge…
…Lead Principal Software Engineer (IC5) to help define and build the next generation of AI networking infrastructure powering large-scale GPU clusters and distributed…
…Partner with Network Engineering, Infrastructure Engineering, and Network Automation teams to transform operational challenges into scalable software solutions. Design and implement resilient, observable, and…
…This role sits within a cross-functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations. The…
…Network Architect to join our Cluster Engineering Team and help shape the front-end datacenter and interconnect fabric for the current and next generations…
…Drive HPC data center networking, including high-speed interconnects, spine-leaf fabrics, and capacity planning for GPU-dense research and engineering infrastructure. Architect and…
…Champion automation by establishing engineering standards, identifying strategic opportunities, and partnering with software engineering teams to deliver scalable automation solutions. Drive systemic improvements that…
…Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development. Make…
…We are looking for an outstanding Senior Software Engineer to work on our security team focused on securing at scale infrastructure, high-performance computing…
…Define and drive full-stack enterprise AI factory baseline architectures across compute, networking, storage, virtualization, orchestration, security, observability, and NVIDIA AI software to deliver…
…networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large…
…bare-metal provisioning and network fabric configuration to automated "one-click" firmware rollouts. Build the "Pit Crew" (Observability): Develop a definitive telemetry and diagnostic…
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and…
…stack — GPU compute, high-speed fabric, and large-scale storage systems — advising on configuration, operational best practices, and incident resolution Own the observability strategy…
…Job Description As a Senior Network Engineer, you’ll design, deploy, and operate the high-performance, low-latency network fabric that underpins our GPU…
…a heterogeneous fleet of GPUs and next-generation accelerators in hundreds of cities worldwide. Working alongside AI/ML engineers, hardware partners, and Cloudflare product…
…GPU partitioning strategies Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents Collaborate with ML engineers…
…Experience with GPU-focused hardware and software (e.g., NVIDIA DGX, CUDA, GPU Operator). Background with RDMA-based fabrics (InfiniBand or RoCE) in HPC…
…will set technical direction, drive observability and reliability standards across the organization, and be the kind of engineer that makes the people around them…
…GPU computing with networking by making sure communication primitives are carefully developed alongside GPU hardware capabilities. Join our team of engineers developing the software…
…engineering, with at least 1 year in a support role for an AI service Strong technical background, with knowledge of AI, ML, GPU technologies…
…In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system…
…BA/BS Degree in Electrical Engineering, Computer Engineering, or related field with equivalent practical experience. GPU Expertise: 5+ years of hardware engineering experience with…