Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Solutions Architect, AI Cluster Performance and Telemetry”. A match may be a passing mention rather than the job itself. Titles only.
49 roles across 55 listings · show every listing · page 2 of 2
…assemble, and deploy AI factories worldwide following NVIDIA's reference concepts. This involves architectural systems, power distribution, cooling systems, integration of telemetry and control…
…of systems monitoring, telemetry, and management tools to improve cluster utilization, reliability, performance and workload insight Build repeatable reference architectures, deployment guides, sizing guidance…
…The AI2NE Org strives to be global leaders in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC), Artificial Intelligence…
…Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints. Participate in architecture and design reviews…
…Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-spectrum visibility that supports the smooth functioning of HPC systems and AI…
…Running perf benchmarks for both training and inference. Collaborate with AE, FAE, and Solution Architect teams on validation for customer issues and technical documentation…
…your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other…
…Able to research, architect and drive complex technical solutions, consisting of multiple technologies, cloud services and AI systems — from early prototype to production at…
…performing solutions for AI, HPC, and cloud-scale systems. You will: Define End-to-End Test Strategy: Own and drive the overall test architecture…
…and future-ready networking solutions. Key Responsibilities – RDMA Fabric Design, Architecture and Scale Lead architecture and design of large-scale RDMA fabrics supporting AI…
…telemetry streams from GPU clusters and operationalize predictive AI models at scale. You will work at the intersection of high-performance data engineering and…
…systems, and AI platform teams to deliver scalable infrastructure solutions. Lead performance analysis, bottleneck identification, and system-wide optimization efforts. Define architecture and technical…
…Collaborate with networking, AI infrastructure, hardware, and cloud platform teams to deliver high-performance solutions. Investigate and resolve complex networking, performance, and reliability issues…
…inference, regression analysis, clustering) to analyze robot performance data, identify failure modes, and uncover patterns that inform model architecture and training strategy decisions. Write…
…products and software for AI infrastructure, including DPUs, SuperNICs, Ethernet networking, accelerated compute, and large-scale AI clusters. Foundation knowledge of high-performance networking…
…Radiant is redefining how AI infrastructure is built. We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure…
…You will improve our CI pipelines to ensure that changes to our clusters are safe, predictable, and automated. Optimize for scale and performance: You…
…Collaborate with internal product and engineering teams to develop “NVIDIA on NVIDIA” reference architectures and best‑practice solutions for large‑scale compute and AI…
…end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise. Perform architectural reviews, OS…
…Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints. Participate in architecture and design reviews…
…Preferred Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines. Background in high-frequency, low-latency systems…
…networking solutions for distributed AI training and inference at scale, with a focus on job completion time, failure resiliency, telemetry, scheduling, and placement. Analyze…
…hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key…
…Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams…