Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Manager, Large Language Model Inference”. A match may be a passing mention rather than the job itself. Titles only.
23 roles
…one software programming language - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and…
…This comprehensive toolkit includes an ML compiler, runtime, and application framework that seamlessly integrates with popular ML frameworks like PyTorch, enabling unparalleled ML inference…
The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs…
…Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters…
…Expertise in deploying large-scale training and inferencing pipeline Experience with pre-training, post-training of transformer-based architectures for language or vision A…
…largest asset managers. Our flagship product, Matrix, delivers industry-leading accuracy, speed, and transparency in AI-driven analysis. It is trusted to help manage…
…building and operating large-scale ML infrastructure. - Design and implement core backend services (e.g., job schedulers, resource managers, autoscalers, model serving layers) with…
…services, such as on-demand + managed Kubernetes and Slurm clusters. This platform serves both our internal StaaS products (inference, fine-tuning) and our external…
…for pretraining and model weights for large-scale inference. Build advanced observability stacks for our customers with automated node lifecycle management for fault-tolerant…
…High performance, large-scale ML systems GPU/Accelerator programming ML framework internals OS internals Language modeling with transformers The annual compensation range for this…
…High performance, large-scale ML systems GPU/Accelerator programming ML framework internals OS internals Language modeling with transformers The annual compensation range for this…
…tuning of a wide variety of ML model families, including massive scale multi-modal large language models like Llama, Qwen, gpt-oss, DeepSeek and…
…ready infrastructure Develop large scale data pipelines to handle advanced language model training requirements Optimize large scale training and inference pipelines for stable and…
…create new site capabilities—with a specific focus on integrating large language models into production-grade services that are reliable, performant, and safe. This…
…Direct experience in developing or deploying large scale GPU based AI applications, like Large Language Model, for training and inference Strong background in building…
…Application examples are Large Language Models (LLMs), Computer Vision, Speech, Recommender Systems, and Multimodal architectures. As a Developer Technology Manager, you'll define and…
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs…
…Diagnose and support Machine Learning and/or Large Language Model deployments, including real-time and batch inference, autoscaling, monitoring, logging, and alerting. Serve as…
…4+ years of experience building large-scale, high-performance backend systems. Strong programming skills in one or more languages (e.g., Python, Go, Rust…
…Recommendations, Personalization, Long-term Reward Modeling, Bandits, Transformers, Large-Scale Language Models, LLM evaluation, RLHF reward modeling/alignment Great interpersonal skills including strong written…
…strong emphasis on training Large Language Models (LLMs). Proven track record of successfully deploying and optimizing LLM models for inference in production environments. In…
…Design and build critical platform components such as service discovery, request routing, load balancing, caching, batching, and traffic management for AI inference workloads. - Reliability…
…performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently…