Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Lead Site Reliability Engineer - Operations Excellence for AI Platforms”. A match may be a passing mention rather than the job itself. Titles only.
681 roles across 833 listings · show every listing · page 2 of 28
…or operational excellence for live services. Experience partnering across teams to deliver infrastructure capabilities with clear ownership and supportability. Experience mentoring engineers, leading design…
…leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations…
…Lead vendor evaluations and manage vendor relationships for technology procurement and delivery. Travel to global sites as needed to support infrastructure builds and major…
…Act as a customer advocate, focusing on service excellence and live site reliability for AI workloads. Research & Innovation: Stay informed on emerging AI infrastructure…
…We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. - Excellence…
…We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer…
…This role sits at the intersection of AI infrastructure engineering, data-center deployment, network architecture, capacity delivery, operational excellence, and organizational transformation.The Principal…
…AI Agents, Large Language Models and intelligent decision engines transform fraud prevention and operational efficiency across one of the world's leading crypto platforms…
…operational excellence Develop deep expertise in ML systems, AI infrastructure, and compute orchestration Ship platform capabilities that enable mission ‑ critical AI workloads for customers…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…AI models on Snapdragon platforms Develop and implement core components of Qualcomm AI Stack runtime framework for inference on resource constrained AI systems, operating…
…site reliability engineers to deliver scalable business solutions. Drive adoption of enterprise-approved AI-assisted engineering practices to improve code quality, operational excellence, troubleshooting…
…As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platform - Cloud Foundational Services SRE organization, you will join our Google Cloud…
…operational noise. Use production data and reliability trends to influence architecture, capacity planning, and engineering priorities. Technical Leadership & Engineering Excellence Provide technical leadership for…
…The Technical Solutions Manager operates at the intersection of business operations, data, and engineering, supporting emerging non-standard scenarios including AI platform partner motions…
…operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for…
…This role is responsible for reducing on-call queries and incidents over time. Qualifications: 7+ years of experience in cloud operations, site reliability engineering…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…Own platform engineering areas including AKS fleet management, monitoring and alerting, CI/CD automation, and production operations. Drive operational excellence by troubleshooting live-site…
…Lead end-to-end delivery and live-site operations of partner-facing services and AI-enabled platforms, ensuring high availability, scalability, and operational excellence…
…Ensure Reliability & Operational Excellence: Maintain live site health, develop incident response playbooks, lead root cause analyses (RCAs), and implement systemic improvements to reduce incident…
…Lead the engineering strategy and execution for enabling customer AI agents to securely and reliably consume Twilio platform capabilities. Build and scale systems that…
…Architect global troubleshooting methodologies and escalation frameworks to enhance the team's operational excellence and engineering standards Co-Lead the team's AI-native…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for…