Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Site Reliability Engineer - Data Center”. A match may be a passing mention rather than the job itself. Titles only.
319 roles across 389 listings · show every listing · page 1 of 13
…2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments. Proven expertise in firmware analysis, hardware specifications…
…ABOUT THE JOB As a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our…
…data center, and AI workloads. We are seeking a Principal Electrical Engineer – Server Hardware Validation to join our OCI Hardware Engineering team On-Site…
…data center hardware support services to ensure resilience, reliability, and availability at scale. Designs, governs, and signs off on advanced change activities and site…
…Analyze electric tariffs and wholesale market prices to evaluate energy cost implications for new and existing data center sites Develop and maintain financial models…
…and firmware ecosystem - Experience with data center power and thermal design - Background in chaos engineering or similar reliability testing methodologies - Understanding of compliance frameworks…
…for Kubernetes services, workloads, and platform reliability. You - 6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge…
PlanetScale is growing rapidly and reinventing the database space. Our database platform offers developers the fastest and most reliable Postgres and Vitess/MySQL databases…
…chain operations and governance across a portfolio of data center campuses, ensuring the safe, reliable, and efficient delivery of infrastructure objectives. The role oversees…
…safety) within the data center. Escalates per applicable policies and standards. Utilizes telemetry, control systems, and other platforms to monitor site status, analyze past…
…landscaping, waste, and pest control—to deliver reliable service with minimal disruption to occupied buildings and site operations. Own day-to-day work orders…
…This is facilities maintenance, not data center facilities operations. You will not own the central utility plant, rack-level cooling, or white-space infrastructure…
…hardware movements, site installations Site sustainment operations: Tier 1 troubleshooting and user engagements, issue resolution and escalation, reporting, working with data Communicate plans and…
…of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage…
…need you as a Data Center Critical Environment Technician Manager. Microsoft’s Cloud Operations & Innovation (CO+I) is the engine that powers our cloud…
…if needed), clear blocking issues, and may enable on site and/or remote data center support activities. Performs assigned tasks and escalates issues during…
…Cloud storage expertise — managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale — is a plus…
…dbt Labs pioneered analytics engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted…
…reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at…
…Your work directly enables AWS to meet datacenter capacity commitments by ensuring manufacturing infrastructure scales efficiently, reliably, and repeatably across a global site network…
…to data center/network sites and Amazon/customer offices as needed. About the team Within AWS Networking, the Backbone Enterprise, and Regional Engineering (BERE…
…To maximize the quality of delivery and reliability of our sites you will: Work across multiple LOB effectively, negotiating and bringing together stakeholders & teams…
…Engineer is responsible for diagnosing, testing, repairing, and validating next-generation AI compute hardware supporting Oracle Cloud Infrastructure (OCI) Go Big data centers. This…
…Ability to connect MCAD model structure, data architecture, automation logic, PLM/BOM requirements, and manufacturing deliverables into a reliable end-to-end process. Strong…
…cloud infrastructure at scale (AWS), observability tooling (Datadog or similar), and deployment pipelines at companies where reliability is a product requirement - Experience owning a…