Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer, GPU Inference”. A match may be a passing mention rather than the job itself. Titles only.
50 roles · page 1 of 2
…Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus. Bachelor’s or Master’s degree in…
…They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference…
…Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions, ensuring every FLOP counts in delivering optimal performance for…
…You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to…
…operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform…
…Our software keeps every machine working reliably hours from the nearest engineer, in harsh environments and on lossy connectivity, doing real work on billion…
…Evaluating, tuning, and maintaining AI/ML models (which includes collecting and preparing data as needed) Evaluation & Inference: Implement algorithms and software to analyse and…
…software team with high standards! This software engineering role involves developing tools for AI researchers and SW/HW teams running AI workload in GPU…
…They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference…
…Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is…
…combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior Backend Engineer, you will play a key…
…Deep understanding of deep learning systems, GPU acceleration, and AI model execution flows. Solid software engineering skills in C++ and/or Python, with strong…
…the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The Together Cloud team builds the Together GPU Clusters product, which…
…About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part…
We are looking for a Storage Software Architect to join the architecture group. You will be part of a team that shapes the next…
…The software architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to…
…time inference, and communication optimization across large-scale systems. Join our world-class team of researchers and engineers building next-generation software and hardware…
…You will be focused on working with engineering to understand the technical capabilities of our inference stack from GPUs, CPUs, networking, CUDA libraries, model…
…We're looking for a Software Engineer focused on Performance Optimization to help push the boundaries of speed and efficiency across our AI infrastructure…
…largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love…
…About the Role We’re looking for a GPU Inference Engineer to contribute to improvements in model serving efficiency for our Robotics research. This…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…The team develops and owns the software stack around NCCL (NVIDIA Collective Communications Library), which enables multi-GPU and multi-node data communication through…