Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Member of Technical Staff - Multi-Modal, Audio”. A match may be a passing mention rather than the job itself. Titles only.
14 roles across 15 listings · show every listing
…This role sits at the center of applied audio model development, working directly with the technical lead to ship production systems that run on…
Overview Microsoft AI is looking for a Member of Technical Staff, Multimodal Infrastructure to help build the next wave of capabilities of our personalized…
…where we are building the next generation of foundation models across different modalities: text, image, video, audio. If you are passionate about exploring, designing…
…and scale across a growing catalog of on-device use cases spanning vision, audio, speech, and multi-modal models. What You'll Do Compiler…
…Nice-to-have: - Audio or speech model experience. - Function calling / tool-use fine-tuning experience. - Multilingual model or data experience. - Background at automotive, autonomous…
…of music through AI-powered visual and audio experiences. What Success Looks Like: • Ship customer impact - Drive top Amazon Music priorities • Build technical excellence…
…audio AI are needed to overcome these challenges and make voice interaction accessible to everyone. THE ROLE As a Member of the Research Staff…
…powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a…
The Amazon Music Customer eXperience Infrastructure (CXI) team delivers the best audio and visual experiences for music, podcasts, and audiobooks across devices and modalities…
…of research experience in large-scale pre-training/mid-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff…
…You will join the multimodal team to push toward superhuman multimodal intelligence. Advance understanding and generation across modalities—image, video, audio, and text—spanning…
…Build and optimize large-scale data pipelines to ingest, process, and analyze multi-modal data (images, video, audio), fueling continuous improvement and personalization of…
…Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality and coherence. Implement efficient algorithms for state-of-the…
…THE OPPORTUNITY Our Data team powers Liquid Foundation Models across pre-training, vision, audio, and emerging modalities. Public data sources are plateauing. Model performance…