Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Solutions Architect, Inference Deployments”. A match may be a passing mention rather than the job itself. Titles only.
644 roles across 752 listings · show every listing · page 2 of 26
…Develop highly reliable, high-throughput inference systems that serve the best AI models internally across SpaceX Architect and implement scalable distributed infrastructure for model…
…Full ownership of your code will be encouraged, including deployment, testing, evaluation, and validation of both traditional and AI components. You will lead a…
…Full ownership of your code will be encouraged, including deployment, testing, evaluation, and validation of both traditional and AI components. You will lead a…
…10+ years’ experience delivering enterprise-scale solutions on Microsoft platforms (Dynamics 365 / M365 Copilot/ Power Platform / Azure), including integration, security, and deployment in customer…
…to ensure customers achieve their AI inference goals on Mantle and Amazon Bedrock Architect production-ready inference solutions leveraging Mantle's distributed engine, Zero…
…AI solutions that can include integration of LLMs/multi-modal FMs in large scale systems, fine-tuning LLMs, deployment and distributed inference of LLMs…
…Develop and maintain device drivers for Linux on ARM and x86 architectures 4. Build automation solutions using modern programming languages (Python, Ruby, Java, C…
…DevOps solutions, including Python, Ansible, Container Runtimes, Kubernetes, and data center deployments Familiarity with AI workloads, including agentic & RAG-based workflows, inference at scale…
…various generative AI model architectures such as LLMs Strong understanding of hardware acceleration and deployment of generative AI inference on edge devices as mobile…
…Job responsibilities Hands-on design, development, and deployment of advanced AI, GenAI, and Agentic business solutions. Partner with engineering leads and architecture to implement…
…and build solutions to enable AI Platform to deliver a global service that handles large scale Microsoft Foundry training and inferencing workloads. As a…
…Optimize and deploy models into production autonomous-driving systems, working across model architecture, inference, and onboard constraints. Set technical direction for geometric vision at…
…AI solutions that can include integration of LLMs/multi-modal FMs in large scale systems, fine-tuning LLMs, deployment and distributed inference of LLMs…
…AI solutions that can include integration of LLMs/multi-modal FMs in large scale systems, fine-tuning LLMs, deployment and distributed inference of LLMs…
…production deployment, meeting customer requirements and delivery timelines. The successful candidate will be equally comfortable debating the merits of deep learning architectures in a…
…inference stack across Amazon and the Open Source Community. As you design and code solutions to help our team drive efficiencies in software architecture…
…scalable ML systems (batch and real-time inference), including data/feature pipelines, model training, evaluation, deployment, monitoring, and drift/performance management Identifies opportunities to…
…About the Role - Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems. - Own the end-to-end…
…of training and inference runs to validate numerical behavior, identify bottlenecks, and guide architectural tradeoffs. You will collaborate closely with architecture, microarchitecture, compiler, runtime…
…In the AI-first GCID organization, Architects are expected to embed AI-native thinking into delivery models, ensuring solutions are intelligent, scalable, and aligned…
…Responsibilities Identify requirements, scope solutions, estimate work, schedule deliverables Apply strong engineering principles for defining robust and maintainable architectures and designs Collaborate broadly across…
…Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution - Experience with AWS solutions such as EC2, DynamoDB, S3, and…
…3) Architecture de solutions ML évolutives et d'opérations ML (MLOps) via les services AWS, en utilisant des solutions GenIA lorsque pertinent. 4) Collaboration…
…We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the…
…Understanding of AI inference serving architecture (vLLM, TensorRT-LLM, Triton) and how network infrastructure supports inference traffic patterns including: prefill/decode disaggregation, KV-Cache…