Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Technical Program Manager - Cluster Orchestration & Applied Training”. A match may be a passing mention rather than the job itself. Titles only.
40 roles across 43 listings · show every listing · page 1 of 2
…CoreWeave is seeking a Technical Program Manager to lead complex, cross-functional programs across Cluster Orchestration and Applied Training within our AI/ML Platform…
…technical discipline, or equivalent experience 7+ years of proven experience as a Solutions Architect, Field Engineer, Infrastructure/Systems Engineer, or Technical Account Manager in…
…Design, implement, and manage automated CI/CD and Continuous Training (CT) pipelines for machine learning model development, evaluation, and delivery. Model Deployment: Containerize, deploy…
…Deep experience with large-scale cluster management systems Strong programming capabilities in Golang and C++ Deep understanding of advanced operating systems, scheduling, containerization, sandboxing…
As a Technical Program Manager for the Platform team, you will partner with engineering teams to directly accelerate the development and maturity of the…
…model training and MLOps tooling in the morning and cluster networking in the afternoon, and who finds the space between infrastructure and application more…
…and deployment. - Establish critical path visibility and aggressively manage schedule compression. Cross-Functional Leadership - Orchestrate execution across: - Real estate & site selection - Power & energy strategy…
…integrations specific to GPU clusters and AI/ML training systems * Manage network partition configurations for multi-node GPU clusters * Handle firmware validation and consistency…
…network repair and remediation programs, ensuring that the high-performance fabrics underpinning Meta's AI training and inference clusters remain operational, resilient, and optimized…
…to cluster schedulers, resource managers, or large‑scale job orchestration systems (e.g., Kubernetes, Slurm, Ray, custom internal systems). Understand modern ML training and…
…You'll have the opportunity to contribute meaningfully to the industry while working alongside talented scientists, engineers, and technical program managers (TPMs) to create…
…Your expertise in orchestration and optimization will be instrumental in advancing our managed Kubernetes and AI training clusters, ensuring they lead the industry in…
…Architecture, Program Management, and field teams to develop product direction, architecture requirements, roadmap inputs, and execution plans Develop and validate scalable cluster designs for…
…Role Responsibilities - Deploy and Manage Kubernetes Clusters, deployed at scale to support AI centric workloads, across both our bare metal clusters and via trusted…
…Own Kubernetes at depth — clusters, networking, operators, container lifecycle, and multi-tenant orchestration. Design, develop, and optimize distributed services and cloud-native infrastructure on…
…We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on…
…Manage Infrastructure-Backed Training Logistics: Oversee the operational deployment of NVIDIA Mission Control (NMC) clusters for hands-on technical training, managing resource contention and…
…Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads. - Cluster Orchestration: Develop…
…Your deep expertise in orchestration and optimization will be instrumental in advancing our managed Kubernetes and AI training clusters, ensuring they set the industry…
…integrations specific to GPU clusters and AI/ML training systems * Manage network partition configurations for multi-node GPU clusters * Handle firmware validation and consistency…
…response, escalation management and root cause analysis in high-pressure, 24×7 production environments. - Stakeholder & Customer Engagement: Act as a trusted technical expert, engaging…
…You will design, build, and optimize large-scale ML and data infrastructure across on-premises NVIDIA DGX clusters and AWS Cloud, enabling advanced training…
…Come help manage digital risk by applying security through identity integration and controls. We are looking for a Senior Architect to join our Identity…
…Holistic Training Services: Beyond Slurm, drive the development of next-generation orchestrators and automated training-based evaluation frameworks that ensure model quality throughout the…
…Manage modern ICT systems at the orchestration layer, including technologies such as OpenShift, Kubernetes, Docker and Terraform Manage modern ICT systems at the application…