Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Manager, Network Reliability and Resiliency”. A match may be a passing mention rather than the job itself. Titles only.
852 roles across 1,103 listings · show every listing · page 2 of 35
…managers, business stakeholders, and other engineers to craft better software. Nice to Have: Fundamental understanding of Linux and all layers of the networking stack…
…Influence reliability and resilience improvements across products and services. Participate in on-call rotations, ensuring timely response and resolution for incidents within assigned service…
…TrafficShift is the critical service that keeps AWS's global infrastructure resilient — enabling engineers to drain and restore traffic across thousands of network devices…
…such as Data Domain , PowerProtect Data Manager (PPDM) , and the broader PowerProtect suite—have the resilient, flexible, and high‑performance software systems required in…
…and resilience management for large-scale, distributed systems, Partner with engineering on public cloud best practices (AWS or equivalent) for compute, storage, networking, messaging…
…Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support. Develops and…
…Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management. Experience optimizing training input pipelines (sharding, prefetching, caching, format choices…
…Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations. Understanding…
…Strong API security fundamentals (OAuth2/OIDC, mTLS, JWT, secrets/cert management) and platform reliability/observability practices. Experience defining API standards and improving developer experience…
…Manage secrets using cloud-native services and Kubernetes patterns, including secure access and rotation practices. Apply security best practices across cloud and Kubernetes (network…
…network, database, and application components. Extensive knowledge “Replication” technologies’ such as IBM CSM and GDPS Expertise in administering z-Series and Hardware Management Console…
…Corporate Network provides best-in-class, scalable, reliable, and resilient network connectivity for data, voice, video and wireless at all corporate office locations globally…
…Implements data governance policies and procedures for data handling (e.g., data retention) to manage data consistency, integrity, accuracy, and reliability throughout the data…
Leads complex, data center hardware support services to ensure resilience, reliability, and availability at scale. Designs, governs, and signs off on advanced change activities…
…Engineers and oversees fault‑tolerant, in‑service‑upgradable designs; optimizes resilience mechanisms (load‑shedding, throttling, rate‑limiting); and sets SLO‑aligned durability and availability…
…Communication, punctual, problem-solving, coaching & development, empathy, operational excellence, stakeholder management. Preferred Qualifications Demonstrated ability to stay resilient and adaptable when navigating shifting priorities…
…We orchestrate and drive deep engagements in areas like Incident Management, Problem Management, Support, Resiliency, and empowering the customers. We represent the customer and…
…and networking needs. - Reliability Engineering: Help design fault tolerant, distributed control plane systems to ensure high availability and resilience as Crusoe Cloud's network…
…Develop and maintain risk management plans to ensure continuity and resilience during deployment and sustainment operations. Document lessons learned and implement best practices in…
Infrastructure Services is part of IS&T and the foundation of Apple's global network operations — managing data center equipment and systems to deliver…
…Serve as the technical advisor to address and resolve critical reliability concerns, advising on and guiding the architecture of highly scalable, and resilient workloads…
…organisation. • You’ll manage the overall technical relationship between AWS and our customers, making recommendations on security, cost, performance, reliability and operational efficiency to…
…efficiently, and reliably. Key job responsibilities - Work on Tier 1 production services that power fulfillment at Amazon scale - Build resiliency, stress testing, and chaos…
…We are looking for a Software Development Engineer (SDEII) to join the management plane and insights team responsible for EKS availability and resiliency, AWS…
…Evolve working early-stage systems into reliable, scalable platforms that support more vehicles, locations, users, and experiments - Own Meaningful Systems: Take services from design…