Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Staff Site Reliability Engineer – Automation and Platform”. A match may be a passing mention rather than the job itself. Titles only.
337 roles across 395 listings · show every listing · page 2 of 14
…automate data movement across the systems, clouds, engines, and tools they rely on. dbt Labs pioneered analytics engineering, helping teams transform data into reliable…
…Designing and implementing machine learning pipelines that support high-performance, reliable, scalable, and secure ML workloads 3. Designing scalable ML solutions and operations (MLOps…
…Experience in site reliability or observability engineering, including monitoring, alerting, service level objectives and indicators, runbooks, and incident response. Experience setting technical direction for…
…Procurement, and external logistics and integration partners on new site buildouts; and with our production networking teams, whose bar for reliability and automation is…
…AI generates the pixels and prototype. You define the patterns, behaviors, and quality bar. - Partner closely with product managers, engineers, and researchers. You shape…
…and architectural decisions, emphasizing scalability, robustness, and reliability — automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance…
…senior and staff levels, and running an effective on-call rotation for a team the rest of engineering depends on. - The infrastructure platform: a…
…and performance engineering for the platform’s critical paths and shared infrastructure. REQUIRED QUALIFICATIONS: 10+ years of experience in site reliability engineering, production engineering…
…of experience as a Site Reliability Engineer, Platform Engineer, DevOps or equivalent infrastructure role. - Can comfortably debug and automate Kubernetes clusters. - Not hesitant to…
…We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This…
…systems, and career sites. About the Role Fivetran + dbt Labs is looking for a Staff Software Engineer to join our dbt Platform Services teams…
…automate data movement across the systems, clouds, engines, and tools they rely on. dbt Labs pioneered analytics engineering, helping teams transform data into reliable…
…automate data movement across the systems, clouds, engines, and tools they rely on. dbt Labs pioneered analytics engineering, helping teams transform data into reliable…
…We are seeking a Senior Staff Software Engineer to build the runtime foundation for NVIDIA’s enterprise AI platforms. You will provide technical leadership…
…Identify and implement optimizations for workloads running on multi‑SoC and multi‑card systems. 3. Site Reliability Engineering (SRE) Apply SRE fundamentals including monitoring…
…Demonstrable hands-on engineering ability: you ship automation or tooling in Python, Go, or similar and can build integrations with AI tools and agentic…
…THE ROLE Architect and own the cloud platform that every engineer at Headway deploys on. Make deploys boring, scaling automatic, infrastructure self-serve, and…
…Track record leading multi-disciplinary teams of senior and staff-level engineers across processes such as SMT, molding/encapsulation, coatings, optics, and automation. Strong…
…and deployment automation for data platforms and pipelines. Strong understanding of platform operations concepts such as observability, monitoring, alerting, incident management, and reliability engineering…
…platform experience. Experience with commercialization and customer-facing programs. Experience building large automation and validation infrastructures. Knowledge of functional safety (FuSa) and reliability engineering…
…and performance engineering for the platform’s critical paths and shared infrastructure. REQUIRED QUALIFICATIONS: 10+ years of experience in site reliability engineering, production engineering…
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google…
…execution engine, infrastructure, data architecture, and frontend platform evolve together as both our product and engineering organization scale. This is a Staff role with…
…8+ years of industry software engineering or site reliability engineering experience. A demonstrated history of reducing operational toil through automation, including transitioning teams from…
…Develops, creates, and modifies software to build large scale, distributed, and fault-tolerant applications, integration between platforms, and microservices. Committed to write reliable, scalable…