Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Senior Software Engineer, Quantized Inference”. A match may be a passing mention rather than the job itself. Titles only.
82 roles across 88 listings · show every listing · page 3 of 4
…This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for development and performance optimization…
…Solid understanding of distributed training systems, scaling laws, and inference optimization techniques. Experience with model optimization methods such as quantization, sparsity, pruning, distillation, and…
…video analytics, software engineering agents, and industry-specific copilots. Advise on inference optimization with TensorRT-LLM, vLLM, and SGLang, including quantization, speculative decoding, and…
…teacher–student, quantization-aware approaches, latency/cost-driven optimization) Robust evaluation and failure-mode analysis for real-world deployment Strong software engineering skills in…
…Or specialized experience in runtime optimizations, model quantization, compression, on-device inference, GPU inference, pytorch, kernel development Your Location: This position is US - Remote…
…Solid understanding of on-device AI / edge AI technology stack , including inference engines, model quantization, and LLM deployment on embedded platforms. Familiarity with mainstream…
…versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference — from novel…
NVIDIA is hiring exceptional software engineers to build and optimize the core inference infrastructure for large language models. Join the TensorRT‑LLM team - the…
…with model optimization for embedded deployment (quantization, pruning, NPU mapping) - Experience with AI accelerators and inference engines (TensorRT, ONNX Runtime, ExecuTorch, etc.) - Knowledge of…
…key workloads with ultra high-speed inference. ABOUT THE ROLE We are hiring a Senior Performance Engineer to join our Product team. You are…
…Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks, and parameter-efficient…
…production constraints Strong software engineering foundations (Python/C++), containerization, AI accelerators, and profiling tools; fluency with modern inference/runtime stacks. Model/system benchmarking and…
…Minimum Qualifications: • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering…
…Minimum Qualifications: • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering…
…Knowledge of Model quantization and compression techiques is a plus Working on ML inference optimizations is a plus Experience working on any AI HW…
…and inference evaluation perspective. · Understanding of ML architectures (Transformers, LSTM, GRU, diffusion models) for validation, benchmarking, and analysis. · Experience validating model quantization, compression, and…
…Collaborate with senior engineers, ML teams, and platform teams to deliver high-quality software. Follow established software engineering practices, including code reviews, documentation, and…
We are looking for a Senior Deep Learning Engineer to help bring Cosmos World Foundation Models from research into efficient, production-grade systems. You…
…Engineering, Software Engineering, Systems Engineering, or related work experience. Responsibilities: • Convert, optimize, and deploy AI models from PyTorch and ONNX frameworks for efficient inference…
…25+ years of industry experience in software engineering. 20+ years of people management experience, including managing managers or senior technical leads and customers. Strong…
…Engineering Group, Engineering Group > Software Engineering General Summary: As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next…
…4+ years of relevant software development experience. Deep understanding of transformer models and inference optimization techniques (e.g., quantization, tensor parallelism, or memory-efficient…
…We are looking for a Senior or Staff level Engineer to work on bleeding-edge AI technology. You will architect high-performance software for…
…Experience with modern DL training frameworks and/or inference engines. Fluency in Python, and solid coding/software-engineering practices A proven track-record in…
…Hands-on experience with modern inference platforms (vLLM, SGLang, Torch, TRT, TRT-LLM). Proficiency with Python and C++. Excellent software engineering fundamentals (source control…