Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Site Reliability Engineer, Inference Infrastructure”. A match may be a passing mention rather than the job itself. Titles only.
41 roles across 43 listings · show every listing · page 1 of 2
…6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role. Deep understanding of…
…highly scalable, reliable, and production-grade systems. Architect High-Performance Inference Systems: Design, optimize, and deploy enterprise-scale LLM serving infrastructures. You will push…
…site, while staying tightly connected to Austin-based engineering and program leadership. MLOps Program Ownership: Own end-to-end execution of ML infrastructure and…
…As a Software Engineering Manager, you will lead the architecture and execution of large-scale ML infrastructure, requiring a deep understanding of Large Language…
…machine learning systems and/or platforms. - Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external model providers to build a platform that is highly reliable, scalable, observable…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external AI providers to ensure our platform remains reliable, scalable, and cost-efficient…
…8+ years of experience in customer facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, ideally supporting large‑scale…
…New Deployments - Own the infrastructure deployment for new sites or site expansions end-to-end: chip vendor and OEM dependencies, architecture updates, cloud foundations…
…for GPU infrastructure - A track record of partnering with researchers or ML engineers, and of making data-informed tradeoffs across reliability, cost, and delivery…
…Manufacturing Systems and Infrastructure (MSI) team is an engineering organisation under the Product Operations org. MSI is responsible for the design, development and maintenance…
…Investigate live site issues and implement and deploy fixes. Participate in an on-call rotation. Drive quality engineering via code reviews and design discussions…
…accepting GPU clusters or HPC infrastructure against defined performance and reliability standards. Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring…
…Experience with AI/ML training and inference infrastructure at scale. Strong background in performance tuning, security/compliance implementation, and hybrid/cloud integration. Demonstrated ability…
…reliability, build speed and test infrastructure. Establish standards for observability, telemetry, mobile performance benchmarking, release quality, and production readiness. Mentor Engineering Managers, Staff Engineers…
…You will partner daily with SRE (Site Reliability Engineering), fleet engineering, and the console and API (Application Programming Interface) teams to bring one coherent…
…engine that powers our cloud services. As a CO+I Critical Environment Technician, you will perform a key role in delivering the core infrastructure…
…workloads using high-performance NVIDIA infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-performance DGX…
…our Infrastructure team on inference infrastructure, model serving and reliability, so the platform stays scalable and cost-efficient across billions of inference requests - Partnering…
…The Mechanical Engineer will serve as the owner's technical representative, partnering closely with colocation providers, developers, EPC contractors, OEMs, and internal infrastructure teams…
…The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing…
…in Computer Science, Engineering, or a related field • You have 5+ years of experience in a DevOps or Site Reliability Engineering role • You're…
…in Computer Science, Engineering, or a related field • You have 2+ years of experience in a DevOps or Site Reliability Engineering role • You're…
…in Computer Science, Engineering, or a related field • You have 2+ years of experience in a DevOps or Site Reliability Engineering role • You're…
…Familiarity with reverse-engineering CAN protocols or developing custom evaluation tools is a significant plus. Infrastructure & Automation: Knowledge of software build systems and the…