Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Deep Learning Architect, LLM Inference”. A match may be a passing mention rather than the job itself. Titles only.
275 roles across 313 listings · show every listing · page 3 of 11
…used to quantize LLMs for accelerating inference on specific GPU architectures Solid understanding of Transformer models and challenges involved in serving large transformer-based…
…on NVIDIA GPU architectures (e.g., HGX platforms, NVLink) using deep learning frameworks (PyTorch, NeMo) and inference engines (vLLM, TensorRT-LLM) - Have 8+ years…
…architecture - Identify and eliminate scaling bottlenecks through load testing, profiling, and architectural optimization, targeting cost-efficiency improvements in compute, storage, and LLM inference spend…
…Experience with Inference deployment and optimization software (ex. vLLM, SGLang, FlashInfer, TensorRT-LLM, Triton, Dynamo, TorchAO, etc.) Demonstrable knowledge of GenAI or machine learning…
…NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep learning—from LLMs and generative AI to recommendation, vision…
…Understanding of AI infrastructure, including inference systems, GPU scheduling, distributed serving, or LLM infrastructure. With highly competitive salaries and a comprehensive benefits package, NVIDIA…
…and inference, all on a large scale! We are seeking a hands-on Solutions Architect with deep expertise in backend infrastructure, inference and cloud…
…a deliberately minimal agent architecture that calls LLM APIs directly (tool loops, multi-step execution, checkpointing), primarily in TypeScript. No prior TypeScript is required…
…Large Language Models (LLMs) Retrieval-Augmented Generation (RAG) AI Agents and Multi-Agent Systems Machine Learning and Deep Learning frameworks NLP and Conversational AI…
…Deep technical fluency. You can read an architecture diagram, understand container and VM runtimes, and have a working knowledge of how AI agents, LLMs…
…ML models (e.g., LLMs, large embedding models). Experience with ML infrastructure, profiling tools, or deep learning inference/serving optimizations. Familiarity with accelerator architectures.
…Deep familiarity with NVIDIA's AI software stack, including CUDA, TensorRT-LLM, NeMo, RAPIDS, Dynamo, Triton, or Isaac. Experience driving the implementation of AI…
…A track record of setting technical strategy through architectural judgment and clear communication via influence, not organizational authority. Distributed Systems Expertise : Deep experience designing…
…learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5. THE OPPORTUNITY Deepgram is looking for a Senior…
…Familiarity with test strategies and ML test harnesses for AI inference accelerators. Working knowledge of deep learning fundamentals, transformer architecture, and LLM inference/serving…
…Experience in building high-performance LLM inference systems using SGLang or vLLM. Publications in top computer architecture, systems, and/or ML conferences. Research Sciences…
…Architect High-Performance Inference Systems: Design, optimize, and deploy enterprise-scale LLM serving infrastructures. You will push the boundaries of throughput and latency. Own…
…designing, training, and scaling deep learning models, with at least 3+ years focused deeply on training Large Language Models (LLMs) or Vision-Language Models…
…Architecture and development of modern inference runtimes and execution stacks, covering frameworks like Llama.cpp, vLLM , PyTorch , WinML , DXCGC, and TensorRT-RTX across LLMs…
…cloud platforms, solution architecture, or startup technical engagement. Strong understanding of generative AI, deep learning, machine learning, model lifecycle, inference serving, model optimization, distributed…
…We are looking for a Senior Systems Engineer to help build that layer. This is a deeply hands-on individual contributor role for someone…
…machine learning services, and frameworks. - Strong track record of working with machine learning systems and/or platforms. - Experience in serving LLMs using inference engines…
…a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based…
…NICE TO HAVE - Experience with AI infrastructure, LLM serving, or machine learning platforms. - Experience working with multiple model providers such as OpenAI, Anthropic, Azure…
…The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models…