Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer - Voice AI (Inference Runtime)”. A match may be a passing mention rather than the job itself. Titles only.
32 roles · page 1 of 2
…owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner…
…Solve technically tests problems that exceed the scope of a generalist Software Engineers, specifically around optimizing Generative AI performance across heterogeneous hardware (CPUs, GPUs…
…About the Role As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack…
…a trusted runtime control plane, a growing experimentation engine, and early warehouse-native capabilities. We are winning lower-maturity buyers at healthy rates. We…
…We're an AI-Native engineering team: AI isn't just a tool we use, it's how we work — across specifications, code, test…
…We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software…
…In-depth knowledge of ML converters/compilers and runtimes, and hardware-accelerated ML inference techniques. Strong understanding of generative AI model architectures and their…
…The work sits at the intersection of distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships. About the Role We…
…system and runtime that allows engineering teams build and safely, reliably run their AI agents. You will also build own AI Agents that solve…
…the runtime across real, varied hardware, and set patterns and standards for the endpoint codebase. Partner with ML engineers to embed the inference path…
…s voice AI runs on — the partners whose chips, servers, accelerators (CPU/GPU/NPU), cloud and inference platforms, on-device and edge runtimes, and…
…Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization. Partner with systems engineers to integrate models into Cloudflare…
COMPANY OVERVIEW Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT…
…different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone. As an Embedded AI Engineer, you…
…Strong hands-on experience with AI/ML software, including generative AI models, inference, model serving, or AI application development. Working knowledge of GPU acceleration…
…As part of the AISW engineering team at Qualcomm, you will own the CI, build, and release infrastructure for the Delegates portfolio — ONNX Runtime…
…Partner with ML engineers to operationalize AI capabilities, building the runtime infrastructure and orchestration systems that power AI agents at scale Data & Storage Solutions…
…Software Development Engineers, Data Engineers, and Business Intelligence Engineers, responsible for critical components of ALICE's portfolio, with a primary focus on LLM Inference…
ABOUT THE TEAM We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for…
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer…
…About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments…
…partnerships, and drive the model optimization and runtime engineering needed to deliver production-quality speech AI on constrained platforms. You will be the technical…
…Strong software engineering foundations (Python/C++), containerization, AI accelerators, and profiling tools; fluency with modern inference/runtime stacks. Customer‑facing experience on launching new…
…observed performance and provide feedback to: - hardware architecture teams - performance modeling teams - system and software engineers. - Debug issues across the stack, including software, runtime…
COMPANY OVERVIEW Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT…