Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Solutions Architect, Inference Deployments”. A match may be a passing mention rather than the job itself. Titles only.
51 roles across 61 listings · show every listing · page 1 of 3
…You will solve complex architecture problems with solutions that are extensible and scale, yet always look for ways to simplify. You will work with…
…team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response…
…Knowledge of ML converters/compilers and runtimes, and hardware-accelerated ML inference techniques. Understanding of Generative AI model architectures and their optimization for on…
…The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration. We are hiring LLM Serving Engineers at multiple levels to…
…The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration. We are hiring an AI Performance Engineers at multiple levels…
…Key Responsibilities · Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators. · Drive graph capture and deployment using…
…Job Description The Qualcomm Cloud AI team is developing software solutions for Inference Acceleration. We are seeking an ambitious, bright and innovative engineer who…
…field Strong background in machine learning including model selection, architecting, training, validation, testing, and deployment Experience in building deep learning neural network for computer…
…field Strong background in machine learning including model selection, architecting, training, validation, testing, and deployment Experience in building deep learning neural network for computer…
…writing modules, CI/CD pipelines, deployment governance, and security reviews AWS cloud architecture : multi-account strategies, landing zones, environment isolation, and cross-account role…
…technical solutions Apply and help refine architectural standards, patterns, and best practices within the infrastructure software team Partner with senior engineers and architects on…
…our RDU (Reconfigurable Dataflow Unit) architecture, combining hardware and software into an integrated platform for large-scale AI deployments. Our products include SambaRack, a…
…Familiarity with real-time inference systems and the architectural considerations of serving ML models in low-latency, high-throughput environments. Prior experience in a…
…Modern Deployment Stack: Experience deploying and operating AI backend systems in production (API services, model inference pipelines, orchestration layers), with demonstrated ownership through reliability…
…8+ years’ experience delivering enterprise-scale solutions on Microsoft platforms (Dynamics 365 / M365 Copilot/ Power Platform / Azure), including integration, security, and deployment in customer…
…inference product designed to deliver high performance at low cost. You’ll provide leadership in the application of new technologies to large scale deployments…
…Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution, or experience in computer architecture - Experience in professional software engineering…
…autonomous driving - Track record of successful production ML deployments - Experience with large-scale distributed environments for ML training and inference - History of impactful first…
…support development, deployment, and operations at scale. Architects, deploys, and operates secure cloud and container-based environments for training and inference, including GPU-intensive…
…Responsibilities Architect, design, and implement high-performance solutions to enhance Search & AI platform. Technically lead the development and scaling of our distributed services, ensuring…
…Collaborate with engineering teams to integrate new GPU architectures into Google Compute Engine (GCE) for rapid workload availability. Oversee the lifecycle of accelerator solutions…
…inference, computer vision, or autonomous systems Experienced with high speed compute infrastructure networking and network topologies Familiarity with edge computing architectures, distributed inference systems…
…architecture conversations, critical-path deployment problems, executive decisions, engineering teams, regional operations, and global process transformation.The successful candidate will help provide solutions to…
…As a Principal Solution Specialist for Core Services, you identify the new market opportunities where compute performance, network fabric architecture, and platform reliability unlock…
…Develop highly reliable, high-throughput inference systems that serve the best AI models internally across SpaceX Architect and implement scalable distributed infrastructure for model…