Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Lead Site Reliability Engineer - Operations Excellence for AI Platforms”. A match may be a passing mention rather than the job itself. Titles only.
681 roles across 833 listings · show every listing · page 1 of 28
…for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the AI Machine Learning and Data platform team…
…24/7 reliability for mission-critical systems. RESPONSIBILITIES Build and lead a high-performing team of data center technicians, systems engineers, and infrastructure specialists…
…We are seeking a Principal Site Reliability Engineering Manager to lead a team responsible for building and operating Substrate services in highly regulated environments…
…The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI…
…The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models…
…AFSIM or similar platforms. Skilled in Statistics and AI/ML Excellent communication and presentation skills Must be willing to travel for test field events…
…leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations…
…Site Reliability Engineering (SRE), or Security Engineering, with at least 3+ years in a senior or lead capacity supporting enterprise applications. Multi-Platform Ecosystem…
…AI systems that meet the highest standards of reliability, security, and operational excellence, including on-premise, disconnected, and air-gapped deployments. As an AI…
…Operations: designing and optimizing our data architecture, ensuring integrations are reliable and scalable, and readying the ecosystem for the next generation of AI-powered…
…engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI…
…leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations…
…Database Excellence stage's engineering leader, you'll build and grow a new engineering cluster in India, keep the long-horizon platform program on…
…Establish and evolve engineering standards for performance, reliability, accessibility, security, testability, operational excellence, and long-term maintainability across the client platform. Develop technical leaders…
…technical leaders and engineering teams. Have at least a high level familiarity with the architecture and operation of LLMs. Have a passion for making…
…technical leaders and engineering teams. Have at least a high level familiarity with the architecture and operation of LLMs. Have a passion for making…
…As a Senior Infrastructure Reliability Engineer you will be proactively driving the reliability risk identification, assessment and mitigation for datacenter infrastructure equipment (Example: Air…
…on Twitch-wide initiatives. - Drive operational excellence: observability, reliability, and operational readiness for critical systems. - Raise the engineering bar through design reviews, code reviews…
…Facilities Manager to lead the safe, reliable, and efficient operation of our aerospace manufacturing facility. This role is responsible for facility infrastructure, building systems…
…Qualifications 10+ years of experience in database engineering, performance engineering, site reliability engineering, or large-scale platform operations. 5+ years of engineering leadership experience…
…Drive engineering excellence through code reviews, automated testing, CI/CD practices, operational rigor, reusable solutions, and continuous improvement. Participate in live-site operations, on…
…engineer on the ML Infrastructure and Platform group, you will help architect and build the foundational infrastructure for AI and machine learning operations at…
…We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. - Excellence…
…operates Klaviyo’s high-scale, event-driven backbone that powers reliable, low-latency async workloads for every product team. As a Senior Platform Engineer…
…Lead technical design for distributed storage, Git repository management, performance, reliability, and scalability problems, using data and benchmarking to guide decisions. Partner with engineers…