Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Solutions Architect, AI Cluster Performance and Telemetry”. A match may be a passing mention rather than the job itself. Titles only.
13 roles
…GPU architectures, and the hardware and software components of large-scale AI clusters. Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…projects) Exposure to microservices architecture and distributed systems, and a desire to learn Familiarity with being on-call and performing operations/SRE tasks or…
…At NVIDIA, we are leading the AI computing revolution, delivering deep learning and high‑performance computing solutions that power many of the world’s…
…third-party risk platforms, cyber risk quantification, vulnerability management, and endpoint and data-security telemetry, with a clear point of view on where AI…
We are looking for a motivated Solutions Architect or Engineer with experience in designing and building Ethernet Networking fabrics for AI Factories for the…
…assemble, and deploy AI factories worldwide following NVIDIA's reference concepts. This involves architectural systems, power distribution, cooling systems, integration of telemetry and control…
…of systems monitoring, telemetry, and management tools to improve cluster utilization, reliability, performance and workload insight Build repeatable reference architectures, deployment guides, sizing guidance…
…The AI2NE Org strives to be global leaders in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC), Artificial Intelligence…
…and evolve systems that support performance analysis, telemetry, and optimization for large-scale GPU- and CPU-based clusters used in AI and high-performance…
…Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints. Participate in architecture and design reviews…