Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Deep Learning Architect, LLM Inference”. A match may be a passing mention rather than the job itself. Titles only.
275 roles across 313 listings · show every listing · page 2 of 11
…inference using open-source models alongside vendor solutions. Our responsibilities are to: - Build and maintain infrastructure to host and serve open-source LLMs reliably…
…As a Staff Software Engineer - Insurance, you will be a senior technical leader across two connected fronts: diagnosing and evolving our service architecture as…
…s deepest technical voice on LLM inference performance — owning optimization strategy, benchmarking rigor, and efficiency at scale. You will work directly with senior engineering…
…Partner closely with Product, Design, Frontend engineers, ML engineers, and Infrastructure teams to build a deeply personalized, multimodal conversational experience. Shape the architecture for…
…Design, develop, and evaluate next-generation AI models and systems, from pre-training and post-training methodologies to novel architectures and scalable inference solutions…
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build…
…deep learning fundamentals, with hands-on experience in diffusion models and generative architectures. Experience in pre-training or refining large language models (LLMs), vision…
NVIDIA is the leading full-stack accelerated computing company, powering the next wave of generative AI, agentic AI, deep learning, data science, cloud-native…
…technology, data science, solution architecture, or developer ecosystems. Strong knowledge of machine learning, deep learning, generative AI, real-time inference, data engineering, MLOps, and…
…Deep expertise with Docker, Kubernetes, container orchestration, and real-time inference services. Modern Generative AI Architecture: Hands-on experience designing and deploying RAG pipelines…
…in feature computation Experience serving LLMs or deep-learning models in production, including GPU capacity planning and inference optimization Prior work supporting internal-developer…
…architecture and technical roadmap for Azure Storage, helping shape the storage foundation behind training, inference, and agentic AI workloads. Partner with engineers, architects, product…
…You must have deep technical experience working with technologies related to large language models including LLM architectures, model evaluation, and fine-tuning techniques. You…
…architecture (design patterns, reliability and scaling) of new and existing systems experience - Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles…
…machine learning or deep learning frameworks for training and inference. Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade…
…inference, uncertainty, experimental design, and modern statistical methods. Expert ML skills — advanced modeling, validation, and production-grade deployment at scale. Deep experience evaluating LLMs…
…leading an engineering team - Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - 7+ years of full…
…Generative AI techniques applied to LLM and Multi-Modal learning (Text, Image, and Video). Knowledge of GPU/CPU architecture and related numerical software. Your…
…Experience in data science and applied machine learning, including classical ML techniques, deep learning models, model evaluation, and production deployment. Disclaimer: Certain U.S…
…open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning…
…You bring the technical curiosity and learning agility to ramp up on any domain – LLMs, distributed systems, device firmware, cloud infrastructure – deeply enough to…
…orchestration, and inference against measurable security outcomes. Take AI capabilities from experimentation to production using strong software engineering and MLOps / LLMOps practices. Build continuous…
…Proven track record of using AI tools daily, with deep expertise in serving architectures, inference providers, agentic frameworks, RAG architectures, and scale evaluation systems…
…healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS) Implement infrastructure security best…
…Your job is to make NVIDIA the obvious place to run inference by turning deep optimization techniques into products that a broad range of…