Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Lead Site Reliability Engineer - Operations Excellence for AI Platforms”. A match may be a passing mention rather than the job itself. Titles only.
681 roles across 833 listings · show every listing · page 4 of 28
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch…
…for collaboration and team meetings. Candidates must be able to reliably commute to our Shoreditch, London office when required. As a Senior Solutions Engineer…
…Microsoft’s Cloud Operations & Innovation (CO+I) is the engine that powers our cloud services. As a CO+I DIAT, you will perform a…
…REQUIREMENTS MUST-HAVES - 🧭 Technical account leadership: You have strong experience in Technical Account Management, Solutions Architecture, Customer Engineering, Site Reliability Engineering, or a…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…NICE TO HAVE - Experience with AI/ML infrastructure, high-throughput inference systems, or model platform reliability. - Experience operating multi-tenant, security-sensitive platforms for…
…market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly…
…hardware platform readiness, fleet health, reliability engineering, lifecycle management, and operational excellence across Microsoft’s global cloud infrastructure. We are looking for a Principal…
…efficiently, reliably, and repeatably across a global site network. Possessing strong leadership, influencing skills, and technical depth are critical elements of success for this…
…platforms providing end-to-end traceability, IoT-based tracking, and AI-assisted anomaly detection across complex multi-tier distribution networks - Loyalty and personalization engines…
…operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for…
…You'll lead the Infrastructure team — one of three teams in Engineering Foundations, alongside Agentic Engineering and Eddy, Headway's internal AI platform — that…
…It drives CorpTech engineering forward through strategic investment in data platforms, software platforms, QA and release excellence, ERP engineering, and AI infrastructure. We build…
…What makes you a good fit: - 8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps or equivalent infrastructure role. - Can comfortably…
…We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This…
…Facilities Engineering Leadership Lead all aspects of facility infrastructure management, including planning, design, construction, maintenance, and operations across multiple states. Ensure safe, reliable, and…
…business, and platform stakeholders to deliver a performant and reliable platform that powers AI for customers globally. This is an on-site role based…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…We're looking for an engineering manager to lead our Metadata team, which owns the metadata layer of the dbt platform: the discovery API…
…leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations…
…team of engineers in Dublin, fostering a culture of operational excellence, proactive problem-solving, and blameless learning. Champion Reliability: Drive a "reliable by design…
…We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. - Excellence…
…Work closely alongside the Maintenance Manager and Maintenance Lead to ensure projects are designed for maintainability and reliability; coordinate project-to-operations handoffs including…