Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Manager, Large Language Model Inference”. A match may be a passing mention rather than the job itself. Titles only.
566 roles across 665 listings · show every listing · page 4 of 23
…privacy-by-design, data minimization, access-control models (RBAC/ABAC), encryption, audit logging, and data lifecycle management. Experience conducting privacy and security reviews, threat…
…world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10…
…in C and C++ in large, production system software Understanding of runtime systems: process/thread models, memory management, IPC/RPC, and resource lifecycle Understanding…
…inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar) - Fluency in Python and proficiency in C++ or another systems language - Excellent…
…tools and techniques in the forecasting space, such as TimeGPT, large language model extensions, causal forecasting, and hybrid approaches. - Drive strategic insight generation by…
…source control management, build processes, testing, and operations experience - Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and…
…Familiarity with frontier AI tooling including large language models (LLMs), RAG (Retrieval-Augmented Generation) systems, agentic AI frameworks, and AI-assisted workflows. Experience with…
…ability to design maintainable codebases and APIs that bridge model inference/orchestration layers with user-facing surfaces. - Experience navigating ambiguous, cross-functional projects that…
…Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval…
…experiments and causal inference analyses to measure the impact of new features and model changes on customer outcomes - Build ROI models and business case…
…deploy machine learning models, including LLM-based systems, retrieval models, and task-specific models - Develop and evaluate models using large-scale, noisy, heterogeneous datasets…
…running experiments on large-scale video, image, and sensor datasets - Collaborate with software engineers to productionize models - optimizing for inference latency, accuracy, and reliability…
…technical direction across an organization - Familiarity with ML workflows (training, inference, model management) and the challenges practitioners face - Master's or PhD in Computer…
…Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses. Track record of applying AI/observability techniques to…
…translate AI ambitions into durable platform designs, so that model training, fine-tuning, and inference rest on consistent, well-governed data and compute foundations…
…Strong leadership, communication, and cross-functional collaboration skills Familiarity with concepts such as retrieval, ranking, personalization, inference, context management, or orchestration, with ability to…
…Safety and governance platforms for AI models and agents Inference, routing, orchestration, and policy enforcement systems Evaluation, red teaming, and monitoring infrastructure for AI…
…runtime to enable and tune large-scale training workloads on the latest Trainium instances. You will bring up model architectures that have never run…
…from large-scale security datasets - Architect distributed computing solutions for model training and inference, leveraging Spark/PySpark and AWS infrastructure - Monitor model performance in…
…Experience in pre-training or refining large language models (LLMs), vision-language models (VLMs), or world foundation models (WFMs). Experience working with large-scale…
…researchers building frontier AI systems, including large language models, multimodal models, reasoning systems, training methods, inference systems, model serving, and scalable AI infrastructure. You…
…Experience with Kubernetes, distributed training, and large-scale inference, including DGX Cloud and Run:ai. Background with foundation models for atomistic simulation (e.g…
…Europe - including large on-prem GPU clusters powering DeepL's research and hundreds of on-prem GPU nodes serving production inference - as well as…
…Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution Our inclusive culture empowers Amazonians…
…world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10…