Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Performance Engineer, Inference Systems”. A match may be a passing mention rather than the job itself. Titles only.
108 roles across 109 listings · show every listing · page 1 of 5
…Shape the core inference backbone that powers Together AI’s frontier models. Solve performance-critical challenges in global request routing, load balancing, and large…
…Staff Infrastructure Engineer to act as a primary technical lead, engineering the 'paved road' for our knowledge retrieval and inference engines. You won't…
…AWS Neuron is the software of Trainium and Inferentia, the AWS Machine Learning chips. Inferentia delivers best-in-class ML inference performance at the…
…Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions, ensuring every FLOP counts in delivering optimal performance for…
…We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors. We are…
…ABOUT THE ROLE As a software engineer at Wispr, you’ll play a crucial role in building the first capable, habit forming voice interface…
…Design, build, and launch to production new features and improvements aimed at unifying common components across the storage systems Dive into performance issues and…
…memory hierarchy, warps, shared memory/register pressure, bandwidth vs compute limits - Proficiency with low-level profiling (Nsight Systems/Compute) and performance methodology - Strong C…
…mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities…
…Evaluate PPA (performance, power, area) for hardware features and system-level architectural trade-offs. Work closely with peer architecture teams and product management to…
…NICE TO HAVE - Experience with AI/ML inference or training infrastructure - Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency…
…a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products. As generative media reshapes…
…We continuously refine our routing engine to calculate more efficient routes, deliver highly accurate ETAs, and manage scalable traffic for every journey, adapting as…
…About the Role We are seeking a versatile and experienced engineer to join our Inference Core Model Bringup team. This team is responsible to…
…an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis…
…the-baseten-inference-stack/ - Driving model performance optimization https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/ RESPONSIBILITIES Core Engineering Responsibilities - Design…
…ML inference and other GPU workloads on embedded compute (e.g., CUDA, TensorRT) - Experience with Rust in production or performance-critical systems - Prior work…
…We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors. We are…
…high-performance commercial engine. You'll own the entire Ray Data product roadmap in a competitive landscape, working closely with the engineering team, the…
…paired with distributed systems infrastructure for long running agentic tasks. 2. The world’s most powerful and scalable LLM inference engine - a distributed, asynchronous…
…paired with distributed systems infrastructure for long running agentic tasks. 2. The world’s most powerful and scalable LLM inference engine - a distributed, asynchronous…
…Scale & Performance: Experience training models across distributed systems (multi-GPU/multi-node) and optimising training and inference performance (e.g., XLA, Triton, CUDA, Pallas…
…scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads…
…Design and implement daemons, frameworks, and system services that orchestrate AI-driven components and manage data, communication, and performance across distributed systems Prototype and…
…ABOUT THE ROLE As a Software Engineer on the Product team, you will build and scale the user-facing systems that sit directly on…