Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Site Reliability Engineer, Inference Infrastructure”. A match may be a passing mention rather than the job itself. Titles only.
193 roles across 213 listings · show every listing · page 3 of 8
…Are you ready to be part of something outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our…
…Engineering, or a related technical field (or equivalent applied experience) Preferred qualifications, capabilities, and skills Site Reliability Engineering (SRE) experience and familiarity with reliability…
…engineering, data science, or software engineering Familiarity with Ray, MLflow, Prefect, or similar distributed compute and workflow orchestration platforms Experience with LLM inference infrastructure…
…complex issues across distributed cloud systems and AI inference infrastructure using logs, monitoring tools, and engineering best practices. - Document test plans, test cases, results…
Amazon Search owns the global search engine for the Amazon shopping site. The Japan team improves the search engine and related systems with new…
…Key job responsibilities As an ML Engineer you will: Lead development of services and infrastructure at the intersection of machine learning, big data, and…
…focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for…
…MDP is scaling from a founding team to 70+ engineers across multiple U.S. sites. You will be joining early, working on hard problems…
…SKILLS & QUALIFICATIONS - 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering. - Hands-on experience building or…
…ABOUT YOU - Strong software engineering skills, particularly in Python, with a track record of building reliable, scalable infrastructure or platform systems. - Experience building sandboxed…
…focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for…
…Engineering Group, Engineering Group > Hardware Engineering General Summary: The Qualcomm Data Center AI System Hardware and Validation Engineering team develops rack-level AI inference…
…infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the large language model inference…
…Infra Reliability · SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and…
…Redwood's sites, lead cross-functional design reviews, mentor junior engineers, and drive continuous improvements in design processes and system reliability. Key Responsibilities Design…
…6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role. Deep understanding of…
…highly scalable, reliable, and production-grade systems. Architect High-Performance Inference Systems: Design, optimize, and deploy enterprise-scale LLM serving infrastructures. You will push…
…site, while staying tightly connected to Austin-based engineering and program leadership. MLOps Program Ownership: Own end-to-end execution of ML infrastructure and…
…As a Software Engineering Manager, you will lead the architecture and execution of large-scale ML infrastructure, requiring a deep understanding of Large Language…
…machine learning systems and/or platforms. - Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external model providers to build a platform that is highly reliable, scalable, observable…
…You'll partner closely with AI Research, Product Engineering, Infrastructure, and external AI providers to ensure our platform remains reliable, scalable, and cost-efficient…
…8+ years of experience in customer facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, ideally supporting large‑scale…
…New Deployments - Own the infrastructure deployment for new sites or site expansions end-to-end: chip vendor and OEM dependencies, architecture updates, cloud foundations…
…for GPU infrastructure - A track record of partnering with researchers or ML engineers, and of making data-informed tradeoffs across reliability, cost, and delivery…