Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Systems Development Engineer, AWS Incident Response (AIR)”. A match may be a passing mention rather than the job itself. Titles only.
35 roles across 40 listings · show every listing · page 1 of 2
…incident response capabilities across AWS regions and services A day in the life A Systems Development Engineer on the AWS Incident Response (AIR) team…
…Provide hands-on architectural leadership in areas like Kubernetes, Serverless, Azure, and AWS. Modernize backend systems with a focus on distributed systems, Go, and…
…Proactively recommend improvements in architecture, deployment, and operations for distributed systems Experience: 8+ years in software site reliability engineering or software development roles. Programming…
…each government agency for which they perform AWS work. Key job responsibilities As a Systems Development Engineer you will: - Design and deliver technology solutions…
…Systems Development Engineers on the team are responsible for maintaining the network tools/systems described above within US GovCloud and other US Government air…
…including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems. Distributed Systems: Hands-on experience developing large-scale distributed systems, databases…
…We are hiring a Machine Learning Engineer to join our ETA team. Our team builds and maintains Lyft's system responsible for estimating/predicting…
…engineering enablement systems, while ensuring compliance with strict government requirements. What you’ll be doing Operate and maintain enterprise grade solutions within air-gapped…
…engineers), AI ambition, and executive air cover to build a genuinely AI-native developer experience: not a bolt-on pilot. This role is responsible…
…building distributed systems that operate at the intersection of consistency, availability, and low-latency replication across AWS regions. Key job responsibilities • Own both the…
…the Software Development Lifecycle Strong understanding of agile methodologies, application resiliency, and security best practices Practical cloud-native engineering experience (AWS, Azure, or GCP…
…Job responsibilities Architects and implements complex, scalable engineering frameworks and solutions using modern software design principles Develops secure, high-quality production code for data…
…including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems. Distributed Systems: Hands-on experience developing large-scale distributed systems, databases…
…White-Glove Incident Response & Escalation Top-tier executive hotline. Carry a monitored cell phone reachable by direct call or text for the most critical…
…Strong understanding of SRE principles, including performance monitoring, incident response, and reliability engineering. Experience with cost optimization strategies for cloud infrastructure. Self-motivated and…
…and implementation work, guides engineers, improves application development practices, and resolves complex production and integration issues across distributed systems. This role will serve as…
…monitoring, alerting, incident response, and continuous improvement of production systems. - Partner cross-functionally with Data Science, Analytics, ML Platform, and Software Engineering teams to…
…RESPONSIBILITIES - Design, build and own backend services for global liquidity management, safeguarding, investment workflows, and cash data management. - Develop systems that ensure funds are…
…systems by pushing for changes that improve reliability and velocity Lead triage and root-cause analysis of high-severity incidents Practice balanced incident response…
…systems reliably: monitoring, on-call, incident response, and reasoning about failure modes up front. - Comfort navigating ambiguity across a range of stakeholders, from engineers…
…Engineers participate in an on-call rotation to maintain high-availability services, leading incident response and driving continuous improvement through deep-dive postmortems. Embedded…
…We are currently looking for an experienced Senior Software Development Engineer to join our team. The ideal candidate is excited about the incredible opportunity…
…incidents end‑to‑end. Infrastructure Skills - Python: Systems tooling and backend services. - PyTorch: LLM Inference engine development and integration, deployment readiness. - Cloud & Automation: AWS…
…usage questions, capacity changes, SLA tracking, incident communications - Build the cross-functional connective tissue between Sales, Engineering, Finance, and customers What We're Looking…
…That means refusing to choose between performance and sustainability, design and engineering, ambition and integrity. In Lucid Air and Lucid Gravity, we have designed…