Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Performance Engineer - LLM Inference Frameworks”. A match may be a passing mention rather than the job itself. Titles only.
164 roles across 181 listings · show every listing · page 1 of 7
…Contributions to inference frameworks such as TensorRT‑LLM, vLLM, SGLang, or similar systems Demonstrated expertise in performance modeling, memory optimization, distributed model execution or…
…Data Science, Machine Learning, AI, LLM, GenAI As a Senior Specialist Solutions Engineer (SSE), ML Engineering, you will be the trusted technical ML expert…
…Optimize model performance for on-device use cases (memory, power, compute constrained environments). Engage directly with research, software engineering, hardware engineering, and product teams…
…Mentor senior and staff engineers across firmware, driver, and framework teams; uphold software quality through architectural and code reviews. Partner with ML and acoustics…
…alongside customer engineering teams during live AI implementation sprints — debugging inference pipelines, optimizing RAG architectures, tuning agent orchestration, or designing evaluation frameworks. You spend…
…Engineering Group, Engineering Group > Software Engineering General Summary: AI Inference Accelerator – Test & Automation - Senior Engineer you will be working for the Qualcomm Cloud AI100…
…Hands-on experience with conversational AI, or LLM fine-tuning and prompt engineering in a production context Exposure to agentic AI frameworks such as…
…data engineers) to create enterprise-scale AI/ML systems that handle high-volume inference workloads, implement comprehensive model and AI governance frameworks, and build…
…Learning Engineer focused on research enablement and performance, turning promising experiments into stable, scalable, user-facing capabilities while making training and inference faster, cheaper…
…major component releases including frameworks, firmware, drivers, and networking components - Establish mechanisms to scale performance engineering practices from current LLM focus to multimodal models…
…deployment workflows Develop reusable ML components, libraries, and frameworks Inference & Performance Optimization Optimize model inference for latency, throughput, and cost Implement advanced techniques such…
…As a senior technical leader, you’ll also mentor engineers, drive best practices, and set the technical vision for AI infrastructure at Pika. What…
…Engineering to spec out logging and with Analytics Engineering to build scalable data tools for mass consumption within Airbnb. Devise business metric frameworks and…
…Production experience with serving frameworks (vLLM, SGLang, TensorRT-LLM, or equivalent), including optimization involving continuous batching, KV-cache strategy, and inference-time quantization. Experience…
We're looking for a Senior Software Development Engineer who wants to build AI-powered products that change how enterprises move to the cloud…
…high-performing team of ML systems engineers. A day in the life You will work with the executive leadership and other senior management and…
…architecture for performance-sensitive systems. Familiarity with modern machine learning and inference system trends, especially around LLMs and generative AI. For senior candidates, strong…
…Research & Develop state-of-the-art Multimodal LLMs and World models to perform 3D Perception using sensor information from Camera, LiDAR and Radar. Integrate…
…LLMs, agentic systems, MCP, RAG, fine-tuning, evals, and responsible AI would be a strong differentiator - Comfortable discussing API integrations, model inference, data pipelines…
…Experience with NVIDIA NeMo (Agent Toolkit, Guardrails, Megatron, Framework, NIM), Nemotron, OSS, Transformer Engine, TensorRT-LLM, Triton, RAPIDS. We are excited to meet researchers…
…senior engineers, raise architectural standards, and influence engineering practices across OCI without requiring direct management authority. Own critical production outcomes, including reliability, performance, security…
…diagnose performance issues, and recommend actions - Lead end-to-end insight development: from data preparation and statistical analysis to LLM prompt engineering that translates…
…monitoring • Drive architecture decisions including model selection, feature engineering from product catalogs, and evaluation metric frameworks • Translate ambiguous, large-scale compliance challenges into well…
…AI systems — retrieval-augmented generation architectures, vector databases, LLM integration, prompt engineering, or agent frameworks- Experience with AWS services (DynamoDB, Lambda, SQS, Bedrock, S3…
…Partnering closely with other engineering teams to translate Safety needs into same and performant software experiences. Performance Optimization: Ensure our features set the benchmark…