Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Site Reliability Engineer - Telemetry”. A match may be a passing mention rather than the job itself. Titles only.
98 roles across 114 listings · show every listing · page 1 of 4
Amazon's Reliability & Maintenance Engineering (RME) organization is looking for a Product Manager - Technical III to own the product lifecycle, roadmap, and stakeholder alignment…
…Our monitoring product gives operators a single, reliable view of device and network health—from live infrastructure telemetry to reliance-aware alarming that points…
…6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role. Deep understanding of…
…data from field/customer, telemetry, and live-site signals. Institutionalizes feedback loops into planning cadences; measures impact. Mentors Staff engineers; partners with Product and…
…Establish technical standards and guardrails that ensure AI-enabled solutions remain secure, reliable, maintainable, and cost-efficient at hyperscale. Use telemetry, experimentation, and data…
…A background in site reliability or observability engineering, including monitoring, alerting, SLOs, SLIs, runbooks, incident handling, and tools such as Prometheus, Grafana, and OpenTelemetry…
…The Global Information Security Office (GISO) at Everpure is seeking an IAM Site Reliability Engineer to operate, automate, and continuously improve our enterprise Identity…
…Proven experience in DevOps, Site Reliability Engineering (SRE), Release Engineering, or a similar productivity-focused discipline. - Testing Infrastructure Knowledge: Deep understanding of automated testing…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external model providers to build a platform that is highly reliable, scalable, observable…
Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external AI providers to ensure our platform remains reliable, scalable, and cost-efficient…
…Knowledge of observability, monitoring, telemetry, and platform reliability practices. #CEAIjobs Data Engineering IC3 - The typical base pay range for this role across the U…
…System design through well-defined interfaces across multiple components, code reviews, leveraging data/telemetry to make decisions. Develop “best-in-class” engineering for our…
…professionals, and using telemetry and other platforms to monitor equipment performance and operations. This role is based 100% on-site in one of our…
…Contribute to Amazon's vision of developing the safest, most secure, reliable, and efficient data centers on Earth. As a Software Development Engineer II…
…As a Lead Site Reliability Engineer at JPMorgan Chase in Infrastructure Platforms, you are a code-first software engineer specializing in reliability. You will…
…As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Data Platform you will solve complex and broad business problems with…
…As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Data Platform you will solve complex and broad business problems with simple…
…As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team , you will solve complex and broad business problems…
…We are seeking a skilled Site Reliability Engineer to design, build, operate, and automate services for traditional IT infrastructure. The ideal candidate will have…
…Thois role is 100% based on-site in one of our Data Centers. Microsoft’s Cloud Operations & Innovation (CO+I) is the engine that…
…providers to enable rapid, reliable deployment of NVIDIA solutions globally. Mentor and guide Infra SA team members and partner engineering teams, sharing best practices…
…Responsibilities Design, implement, test, and operate features for the billing, cost management, and FinOps platform and experiences, with guidance from senior engineers. Build reliable…
…and statistical acceptance criteria that produce defensible reliability evidence. • Correlate bench and accelerated tests with field telemetry, returned hardware, service records, and observed environmental…
…As a Senior Site Reliability Engineer, you’ll work to: Build and scale our internal platform offerings (compute, storage and networking services) to ensure…