Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Deep Learning Architect, LLM Inference”. A match may be a passing mention rather than the job itself. Titles only.
47 roles across 48 listings · show every listing · page 1 of 2
…vLLM, Ollama, LLMLite, TensorRT-LLM, Triton Inference Server, or similar technologies, including the performance and cost trade-offs associated with serving LLMs at scale…
…Driving algorithmic and modeling improvements to the system using primarily deep learning techniques from NLP and computer vision, including the latest LLM models. Deploying…
…We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and…
…AI & Deep Learning Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems. Understand ML inference concepts including batching…
…AI/ML conferences and deep expertise in generative AI, including Multi-Modal Foundation Models, Efficient Architectures, LLM Reasoning, Reinforcement Learning, Agentic AI, and Autonomy…
…inference using open-source models alongside vendor solutions. Our responsibilities are to: - Build and maintain infrastructure to host and serve open-source LLMs reliably…
…As a Staff Software Engineer - Insurance, you will be a senior technical leader across two connected fronts: diagnosing and evolving our service architecture as…
…s deepest technical voice on LLM inference performance — owning optimization strategy, benchmarking rigor, and efficiency at scale. You will work directly with senior engineering…
…Partner closely with Product, Design, Frontend engineers, ML engineers, and Infrastructure teams to build a deeply personalized, multimodal conversational experience. Shape the architecture for…
…Design, develop, and evaluate next-generation AI models and systems, from pre-training and post-training methodologies to novel architectures and scalable inference solutions…
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build…
…deep learning fundamentals, with hands-on experience in diffusion models and generative architectures. Experience in pre-training or refining large language models (LLMs), vision…
NVIDIA is the leading full-stack accelerated computing company, powering the next wave of generative AI, agentic AI, deep learning, data science, cloud-native…
…technology, data science, solution architecture, or developer ecosystems. Strong knowledge of machine learning, deep learning, generative AI, real-time inference, data engineering, MLOps, and…
…Deep expertise with Docker, Kubernetes, container orchestration, and real-time inference services. Modern Generative AI Architecture: Hands-on experience designing and deploying RAG pipelines…
…in feature computation Experience serving LLMs or deep-learning models in production, including GPU capacity planning and inference optimization Prior work supporting internal-developer…
…architecture and technical roadmap for Azure Storage, helping shape the storage foundation behind training, inference, and agentic AI workloads. Partner with engineers, architects, product…
…You must have deep technical experience working with technologies related to large language models including LLM architectures, model evaluation, and fine-tuning techniques. You…
…architecture (design patterns, reliability and scaling) of new and existing systems experience - Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles…
…machine learning or deep learning frameworks for training and inference. Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade…
…inference, uncertainty, experimental design, and modern statistical methods. Expert ML skills — advanced modeling, validation, and production-grade deployment at scale. Deep experience evaluating LLMs…
…leading an engineering team - Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - 7+ years of full…
…Generative AI techniques applied to LLM and Multi-Modal learning (Text, Image, and Video). Knowledge of GPU/CPU architecture and related numerical software. Your…
…Experience in data science and applied machine learning, including classical ML techniques, deep learning models, model evaluation, and production deployment. Disclaimer: Certain U.S…
…open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning…