Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Technical Lead, Cluster Management System”. A match may be a passing mention rather than the job itself. Titles only.
35 roles across 42 listings · show every listing · page 1 of 2
…a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full…
…a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full…
…As a system software architect, lead technical innovation and strategic collaborations with major hyperscalers to architect next-generation data center products. Align NVIDIA's…
…in troubleshooting complex systems issues independently using observability tools and service logs Experience developing and managing highly-available distributed systems Passion for designing thoughtful…
…Our systems integration engineers internalize the nuances of each deployment, ensuring the technical success of the end-to-end solutions we ship. ABOUT THE…
…Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters…
…Knowledge of cloud and cluster level deployment and management systems. Participation and contributions in standards bodies such as OCP and DMTF. Familiarity with CXL…
…REQUIRED QUALIFICATIONS - 4+ years of product management experience with technical products - Strong technical background in distributed systems, ML infrastructure, or data processing - Experience working…
…cluster level deployment and management systems. Experience with GPU computing (CUDA), deep learning workloads. Knowledge of Memory fabric and CXL architectures. NVIDIA is leading…
…a Kubernetes-based PaaS spanning hundreds of production clusters Apollo: secure, fleet-wide deployment and change-management for complex microservice suites Signals: our full…
…a Kubernetes-based PaaS spanning hundreds of production clusters Apollo: secure, fleet-wide deployment and change-management for complex microservice suites Signals: our full…
…Key Responsibilities - Architect and implement low-level control-plane software responsible for system bring-up, configuration, and management of cluster-scale AI compute deployments…
…Your expertise in orchestration and optimization will be instrumental in advancing our managed Kubernetes and AI training clusters, ensuring they lead the industry in…
…Work on a distributed GPU scheduling system for the on-demand clusters product, Instant Clusters. Build out a global management plane for managing our…
…effectively with both technical and non-technical team members Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O…
…Knowledge of cloud and cluster level deployment and management systems Expertise in Out of Band and In-band management architectures. Experience with GPU computing…
…a Solutions Architect, engineer, researcher or technical account manager in cloud infrastructure focusing on building distributed systems or HPC/cloud services, with an expertise…
…Strong experience with Docker, Terraform and at least one major Cloud Platform Strong hands-on experience managing Kubernetes clusters Clear understanding of Infrastructure as…
…Maintain large scale HPC/AI clusters with monitoring, logging and alerting Manage Linux job/workload schedulers and orchestration tools. Develop and maintain continuous integration…
…Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute…
…We're seeking a dedicated technical leader to spearhead the Materialized Views engineering team. The team is responsible for building next generation Materialized View…
…necessary. - Strong technical background in distributed systems software development (K8s and its ecosystem) is preferred. - Technical experience with bare metal cluster management software and…
…Senior Staff Engineering Manager. You will: Lead the design, development and deployment of cutting-edge 4D world models and generative systems for ultra-realistic…
…8+ years of proven experience building and scaling predictive models across distributed systems (eg: Spark, Kubernetes, GPU clusters), production model hosting, and handling end…
…their applications and systems and stay abreast of the latest developments in language models and generative AI technologies. Provide technical leadership and guidance on…