Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “AI Researcher, On-Device LLM Efficiency”. A match may be a passing mention rather than the job itself. Titles only.
29 roles across 33 listings · show every listing · page 1 of 2
…to deliver state-of-the-art multimodal AI models that enable new experiences on Amazon devices. - 3+ years of building machine learning models for…
…Amazon Music customers on Alexa/Echo, mobile, and web. Key job responsibilities - Use machine learning, deep learning, LLMs and Agentic AI techniques to create…
…The LocalAI team is seeking a Senior Systems Software Engineer to build efficient on-device AI software for RTX and DGX-class systems. This…
…The LocalAI team is seeking a Senior Systems Software Engineer to develop efficient on-device AI software for RTX and DGX-class systems. The…
…The Local AI team is seeking a System Software Man a g e r to lead development of an efficient on-device AI software…
…A major focus of the role will be the co-development of frontier models and efficient models for Apple silicon and on-device intelligence…
…on-device Agentic AI technologies that provide intelligent, personalized, and context-aware assistance while preserving efficiency and privacy on edge devices. Our research focuses…
About Zscaler Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an AI-forward enterprise , we…
…of sophisticated on-device and server-side software frameworks that enable low-latency, privacy-preserving, and cost-efficient retrieval, orchestration, and LLM inference across…
…Our team in Seoul is developing AI technology that enables edge devices to run various models—including LLMs, VLMs, omni models, and more—efficiently…
…efficient LLM inference techniques and recent research in this area. - Experience with model quantization or pruning, and familiarity with quantization toolkits (e.g., AIMET…
…The compute performance and power efficiency requirements of our AI devices require custom silicon. Reality Labs Silicon team is advancing research and development in…
…LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML). On-device or edge…
…Drive the strategy, roadmap, and execution of NVIDIA’s inference frameworks engineering, focusing on Client AI. Partner with internal compiler, libraries, and research teams…
As Alexa Audio, we own the audio experiences on Alexa enabled devices. These experiences include Music, Podcast, Audio Books, Radio, and Ambient soundscapes. Our…
…performance, efficiency, and scalability, ensuring a seamless user experience under strict on-device constraints. Develop on-device software that bridges multimodal AI models and…
…of translating research advances into production AI systems Experience with large-scale Omni LLM training optimization, distributed training frameworks, or inference efficiency techniques such…
…particular emphasis on Large Language Models (LLMs) and Natural Language Processing (NLP) Proven ability to understand, interpret, and apply cutting-edge research to consumer…
…to new menus - Push the very boundaries of LLMs and voice AI technology to solve one of the technology industry’s historically elusive challenges…
…to new menus - Push the very boundaries of LLMs and voice AI technology to solve one of the technology industry’s historically elusive challenges…
…EXAMPLE PROJECTS These are some examples of projects that engineers on our team have worked on recently: - Design and build AI agents for large…
…EXAMPLE PROJECTS These are some examples of projects that engineers on our team have worked on recently: - Design and build AI agents for large…
…that AI researchers and post-training teams depend on. The team spans the full software stack, from collaborating closely with the researchers and labs…
…real-time inference efficiently enough to support deployment to all Roblox players. ML Platform Team : The Foundation AI Group is on a mission to…
…to deliver state-of-the-art multimodal AI models that enable new experiences on Amazon devices. - PhD, or a Master's degree and experience…