Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Deep Learning Architect, LLM Inference”. A match may be a passing mention rather than the job itself. Titles only.
274 roles across 312 listings · show every listing · page 5 of 11
…workloads, including LLMs and GPU-based services. You’ll collaborate closely with cross-functional teams - including platform, backend, and machine learning engineers - to design…
…and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5. Deepgram is seeking a Senior Technical Program…
…AWS Specialist Solutions Architects (SSAs) are technologists with deep domain-specific expertise, able to address advanced concepts and feature designs. As part of the…
…AWS Partner Solutions Architects are technologists with deep domain-specific expertise who operate at the intersection of partner strategy and technical depth. As part…
…You will contribute to the efficient execution of advanced deep neural networks (DNNs), large language models (LLMs), and other modern AI architectures. You will…
…understanding of inference workloads, model serving architectures, Kubernetes, containers, APIs, GPU infrastructure, and AI application deployment patterns. - Familiarity with LLMs, RAG architectures, agentic applications…
…This role requires deep technical judgment, the ability to influence senior leaders (Directors, VPs) across Amazon Payments, Amazon Business, and Seller Partner Services, and…
…helping them integrate LaunchDarkly into inference-time and workflow-level decisions. Apply a consultative approach to understand customer architectures, gather detailed requirements, and influence…
…Experience using LLMs to accelerate data engineering (automated data quality checks, intelligent ETL generation, Text2SQL) is a plus. Basic ML/deep learning knowledge and…
…Bonus: - Background in recommender systems, representation learning, or self-supervised learning; production LLM application development; open-source contributions or publications in AI. WHY WEALTHSIMPLE…
…Deep conceptual and practical understanding of deep learning frameworks (e.g., JAX, PyTorch, TensorFlow), large-scale cloud training or inference infrastructure, and LLM ecosystems…
…K nowledge of machine learning, deep learning, quantization, and optimization. Experience with different Neural Network architectures: DNNs, CNNs, RNNs/LSTMs, GANs, LLMs, etc. K…
…architecture and implementation of lighthouse ML and LLM-powered solutions Design and implement highly scalable and reliable data processing pipelines and deploy model inference…
…Deep hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…and/or Windows) - Fundamentals of Machine learning and LLMs, their architecture along with work experience on certain LLM models. Amazon is an equal opportunity…
…experienced Senior Delivery Consultant - Modernization with deep expertise in Artificial Intelligence to join AWS Professional Services (ProServe). This role combines strategic architectural vision with…
…Understanding and working knowledge with Deep Learning large scale training backend like Pytorch, Nemotron and Inference backend like vLLM, SGLang. Has working knowledge of…
…Profiles should be comfortable in a dynamic environment with experience in Deep Learning, LLMs, and GPU technologies. This role is an excellent opportunity to…
…Whether your strength is working with researchers to understand and address their need optimizing deep learning frameworks, or building distributed infrastructure, we want to…
…Profiles should be comfortable in a dynamic environment with experience in Deep Learning, LLMs, and GPU technologies. This role is an excellent opportunity to…
…used to quantize LLMs for accelerating inference on specific GPU architectures Experience in systems engineering fundamentals: caching, CUDA, autoscaling, high throughput, low latency, x…
…Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks, and parameter-efficient fine-tuning (LoRA, QLoRA). Architect and implement…
…model training speed for novel, sophisticated deep learning architectures. Optimize Real-Time Inference: Engineer high-performance model inference solutions to support the seamless deployment…
…The ASEAN Technology Team is seeking a hands-on Senior Solution Architect, AI Engineering — a forward-deployed technical leader who converts AI ambition into…