Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Staff Software Engineer, GPU Inference”. A match may be a passing mention rather than the job itself. Titles only.
208 roles across 227 listings · show every listing · page 1 of 9
…We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working across our custom inference APIs, the vLLM serving runtime…
…inference scaling across multi-node clusters using Ray Serve and Triton Experience in leading technical projects and supporting architectural decisions with data Software Engineering…
…Solve technically tests problems that exceed the scope of a generalist Software Engineers, specifically around optimizing Generative AI performance across heterogeneous hardware (CPUs, GPUs…
…About the team The team is comprise of both Hardware Design Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with…
…design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another…
…We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls…
…You will lead a team of software and AI engineers, coaching and mentoring them. Familiarity with CI/CD, test-driven development, and evaluation-driven…
…We are seeking a Systems Development Engineer to develop automation software, diagnostic tooling, and fleet health infrastructure for our accelerated (AI/ML) server platforms…
…NPU and GPU solutions, enabling optimized inference across diverse hardware platforms. You will have the opportunity to demonstrate your passion for software design and…
…in computer science, electrical engineering, machine learning, distributed systems, semiconductors, or a related field. - Writing about AI models, inference, GPUs, ASICs, datacenters, networking, or…
…Prometheus, Grafana, OpenTelemetry, and alerting people actually act on. - GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training. - AI-assisted operations: Building…
We are looking for a Senior Inference Engineer to own inference for real-time multimodal conversational AI. This is a full-stack inference role…
…This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance model inference solutions for…
…Knowledge of GPU architecture, CUDA, Triton, custom kernels, or hardware-aware optimization. Familiarity with distributed training and model parallelism. Experience with AWS Trainium, Inferentia…
…mentor other engineers. Ways To Stand Out From The Crowd: Experience building platforms for AI/ML training, inference, model serving, GPU-accelerated workloads, distributed…
…As a Software Engineer II on the team, you will help design, build, and harden a secure, evaluable, vendor-agnostic agent platform that teams…
…1) AI Infrastructure with purpose-built chips like AWS Trainium and Inferentia, GPU-powered instances, and optimized frameworks like PyTorch and JAX, 2) AI…
…are optimized for Training and Inferencing. Lead the development of end-to-end AI network architecture including gpu-gpu, server – storage, in-rack cabling…
…Experience with inference-serving frameworks, GPU-aware scheduling, or model-performance optimization. Expertise in vector databases, GPU-accelerated query engines, or distributed data platforms…
…Engineering Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads. Build, deploy, and operate components supporting LLM inference, agentic…
…Proven ability to design, implement, and optimize scalable ML architectures, from distributed training to real-time inference. Strong software engineering skills in Python, C…
…LLM inference, an inter-agent message bus, parametric CAD pipelines, a print farm, and the infrastructure connecting it all. We need an engineer who…
…key workloads with ultra high-speed inference. Key Responsibilities - Work with architects, designers, post silicon and software engineers to ensure a high-quality design…