Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Technical Program Manager, ML Fleet Capacity, Systems Enablement”. A match may be a passing mention rather than the job itself. Titles only.
27 roles · group by role · page 1 of 2
…Lead complex, cross-functional programs related to ML Fleet capacity management, including the design, update, and maintenance of ML Fleet's cluster-level allocation…
…results. - Fleet Foundation — Builds the host enablement tooling systems: OS and ZTP switch provisioning, firmware management, out-of-band access, and power management. The…
…fleets). - Experience with capacity planning, procurement, resource management, or efficiency work on systems built on large private clouds or public cloud providers. - Deep system…
…Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible…
…Experience with large capacity fleets, AI/ML infrastructure, or large-scale inference or training systems. Experience with capacity planning, fleet management, or supply/demand…
…You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management…
…Product Manager to own the product vision and strategy for M365 Fleet Infrastructure, including Data Availability, Durability, Reliability, Data Protection, Resource Efficiency, Capacity Modeling…
…and the systematic use of data and automation to prevent disruptions and drive continuous improvement. We are seeking a Technical Program Manager to own…
…data center or network capacity planning, data center or network infrastructure program/project management, or related technical fields - Experience owning program strategy, end to…
…QUALIFICATIONS: * 10+ years of experience in delivering and managing highly scalable and highly available distributed systems. * Strong knowledge of JAVA and object oriented programming…
…As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people…
…operate the software systems that enable the NPD Interconnects team to monitor, qualify, and manage interconnect products across the AWS fleet. You will build…
…Experience running GPU compute fleets for ML serving and training (e.g., NVIDIA L40S/H100, SageMaker) and associated capacity and cost governance and experience…
…orchestration systems, capacity management, and reliability automation. - Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self…
…runtime systems (model serving, inference optimization, containerized workload execution, or real-time ML pipelines) in a customer-facing or deal-shaping capacity. Deep working…
…Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible…
…Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible…
…The systems you build directly determine how quickly and reliably new hardware turns into usable customer capacity. What You’ll Do Fleet Engineering spans…
…networking, storage systems, operating systems and hands-on systems engineering experience - Knowledge of systems engineering fundamentals (networking, storage, operating systems) - Experience programming with at…
…Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible…
…Build monitoring, alerting, and self-healing automation for the bench fleet. Proactively identify systemic risks — capacity bottlenecks, hardware degradation patterns, infrastructure single points of…
…Build monitoring, alerting, and self-healing automation for the bench fleet. Proactively identify systemic risks — capacity bottlenecks, hardware degradation patterns, infrastructure single points of…
…a Member of Technical Staff – Capacity & Efficiency Infrastructure , to help us improve manage, and improve the efficiency of, our compute fleet. We’re seeking…
…designing, delivering and operating AWS cloud offerings that enable high performance and scalability in AI/ML and HPC workloads. You are intrigued by the…
…Experience with ATS and HRIS systems, such as GreenHouse and WorkDay. Bonus Points Experience managing large-scale operations or volume hiring programs alongside technical…