Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Staff Software Engineer, GPU Inference”. A match may be a passing mention rather than the job itself. Titles only.
13 roles
…They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference…
…You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that…
…We're looking for Senior and Staff Software Engineers to architect, build, and own major areas of our onboard platform, from the OS, drivers…
…Research Engineering (Machine Learning), London We are looking for Research Engineers with different levels of experience - Mid through to Senior, Staff, Principal or equivalent…
…Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is…
…We're looking for a Software Engineer focused on Performance Optimization to help push the boundaries of speed and efficiency across our AI infrastructure…
…The Role We are seeking a Staff Software Engineer to lead the integration and deployment of advanced perception models into Lucid’s production ADAS…
…Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). Strong background in system optimizations: batching, caching, load balancing, parallelism. Low-level inference…
…About the Role We're hiring a Staff Engineer to own major areas of the architecture of our Inference Cloud Platform. This team owns…
…Have significant software engineering or machine learning experience, particularly at supercomputing scale Are results-oriented, with a bias towards flexibility and impact Pick up…
…ABOUT THE ROLE As an engineer on the Supercomputing Platform & Infrastructure team, you will design, build, and operate the large-scale GPU infrastructure that…
…Large‑scale inference systems (e.g., SGLang, vLLM, FasterTransformer, TensorRT, custom engines, or similar), GPU performance, distributed serving. RL‑first profile: RL / post‑training…
…Collaborate closely with x-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions…