Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Solutions Architect, AI Cluster Performance and Telemetry”. A match may be a passing mention rather than the job itself. Titles only.
49 roles across 55 listings · show every listing · page 1 of 2
We are looking for a Senior Solutions Architect specializing in Data Center Systems & Performance to join our elite solutions architecture team. In this role…
…heat exchangers, and thermal management solutions — that complement and unlock the full performance of the PowerEdge server portfolio, particularly for AI and high-density…
…The Senior Network Production Operations Engineer will implement and operate the global edge, backbone, and data center network for high-performance compute (HPC) clusters…
…You will contribute to our cluster observability platform for Trainium accelerators — building real-time telemetry collection, device health tracking, cross-node error correlation, and…
…Architect AI-optimized storage and data management solutions integrating optimized, scale-out file, object, and parallel filesystem SDS technologies.‑optimized storage and data management…
…What you'll do Performance Insights & Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency…
…knowledge graphs, unsupervised learning (clustering, dimensionality reduction, anomaly detection), NLP, and embedding-based retrieval - - Experience designing evaluation frameworks for AI systems where ground truth…
…Deep understanding of enterprise AI cluster operations, multi-node GPU scaling, telemetry/observability tools, and high-performance interconnects (e.g., InfiniBand, NVLink). Experience building…
…Experience with telemetry, observability, network monitoring, and performance analysis. Knowledge of network security principles and high-availability architectures. Excellent troubleshooting, communication, collaboration, and organizational…
…The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and…
…Key Responsibilities The AI2NO Org strives to be global leaders in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC…
…tooling, telemetry, security, and incident response. Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic…
…We are looking for a Senior Solutions Architect who is both customer‑facing and deeply hands‑on with SONiC‑based networking and GPU server…
…and hyperscale AI cluster architectures. Experience developing network stress tools, validation frameworks, performance benchmarks, or observability solutions. Knowledge of packet analysis tools, telemetry infrastructure…
…solutions for search, security, and observability help organizations deliver on the promise of AI. What is The Role : Join the I nfoSec - Security Architecture…
…and evolve systems that support performance analysis, telemetry, and optimization for large-scale GPU- and CPU-based clusters used in AI and high-performance…
…Qualcomm Software Engineers collaborate with systems, hardware, architecture, test engineers, and other teams to design system-level software solutions and obtain information on performance…
…GPU architectures, and the hardware and software components of large-scale AI clusters. Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…At NVIDIA, we are leading the AI computing revolution, delivering deep learning and high‑performance computing solutions that power many of the world’s…
…third-party risk platforms, cyber risk quantification, vulnerability management, and endpoint and data-security telemetry, with a clear point of view on where AI…
We are looking for a motivated Solutions Architect or Engineer with experience in designing and building Ethernet Networking fabrics for AI Factories for the…