Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Solutions Architect, AI Cluster Performance and Telemetry”. A match may be a passing mention rather than the job itself. Titles only.
13 roles
…Architect AI-optimized storage and data management solutions integrating optimized, scale-out file, object, and parallel filesystem SDS technologies.‑optimized storage and data management…
…What you'll do Performance Insights & Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency…
…knowledge graphs, unsupervised learning (clustering, dimensionality reduction, anomaly detection), NLP, and embedding-based retrieval - - Experience designing evaluation frameworks for AI systems where ground truth…
…Deep understanding of enterprise AI cluster operations, multi-node GPU scaling, telemetry/observability tools, and high-performance interconnects (e.g., InfiniBand, NVLink). Experience building…
…Experience with telemetry, observability, network monitoring, and performance analysis. Knowledge of network security principles and high-availability architectures. Excellent troubleshooting, communication, collaboration, and organizational…
…The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and…
…Key Responsibilities The AI2NO Org strives to be global leaders in the RDMA cluster networking domain and enable seamless, accelerated High-Performance Compute (HPC…
…tooling, telemetry, security, and incident response. Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic…
…We are looking for a Senior Solutions Architect who is both customer‑facing and deeply hands‑on with SONiC‑based networking and GPU server…
…and hyperscale AI cluster architectures. Experience developing network stress tools, validation frameworks, performance benchmarks, or observability solutions. Knowledge of packet analysis tools, telemetry infrastructure…
…solutions for search, security, and observability help organizations deliver on the promise of AI. What is The Role : Join the I nfoSec - Security Architecture…
…and evolve systems that support performance analysis, telemetry, and optimization for large-scale GPU- and CPU-based clusters used in AI and high-performance…
…Qualcomm Software Engineers collaborate with systems, hardware, architecture, test engineers, and other teams to design system-level software solutions and obtain information on performance…