Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Solutions Architect, Inference Deployments”. A match may be a passing mention rather than the job itself. Titles only.
26 roles · page 1 of 2
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer…
…You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance…
…Serve as a technical advisor and problem solver with partner engineering teams, collaborating on architecture, code, and integration for Omniverse and AI enabled-solutions…
…architectures quickly and efficiently, collaborating with other teams to determine data and infrastructure support needs, and working to improve model optimization and inference speeds…
…Modern Architectures: Hands-on experience building and working with modern model architectures (e.g., Transformers, GNNs, Diffusion Models). Full ML Lifecycle: Experience taking models…
ABOUT THE TEAM OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building…
…A demonstrated history of reducing operational toil through automation, including transitioning teams from manual deployment processes to self-serve pipelines. Familiarity with LLM inference…
…KEY RESPONSIBILITIES: - Architect and build scalable, resilient, and high-performance backend infrastructure to support distributed training, inference, and data processing pipelines. - Lead technical design…
NVIDIA is seeking a sharp, innovative, and hands-on Architect to help shape the future of LLM inference at scale. Join our dynamic E2E…
…Collaborate directly with the GTM team (Account Executives and Solutions Architects) to ensure smooth integration and successful deployment of ML solutions. - Demo / Proof of…
…understand the technical capabilities of our inference stack from GPUs, CPUs, networking, CUDA libraries, model architectures and deployment techniques (parallelisms, configurations, etc.). You will…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…deployment on memory constrained platforms - Collaborate closely with ML engineers and software developers on technical efforts to find and optimize efficient model architecture solutions…
…Architect scalable distributed systems for data processing, inference, and orchestration of large-scale GenAI workloads. Optimize backend performance for latency, throughput, and cost—ensuring…
…WHAT YOU'LL DO Own the end-to-end architecture of AI Factory tactical data centers and deployable command and control centers Conduct architectural…
…Key responsibilities: - Scope, architect, and deliver innovative high-quality solutions - Effectively communicate complex features and systems in detail - Collaborate with teams across Apple on…
…Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. - Build tools to give us…
…Performance Optimization Optimize inference pipelines using CUDA, TensorRT, and mixed-precision techniques for real-time performance. Implement multithreaded scheduling and containerized deployments for automotive…
…Partner with Solutions Architecture, Research, and Engineering to design solutions/POCs and prove value across inference and post-training engagements. Inform product needs by…
About the Role As a Solutions Architect at Together AI, you will work with customers and prospects to create business value through Generative AI…
…training/inference optimization, integration with cloud-native services, MLOps, etc. Serve as a trusted practitioner for enterprise GenAI solutions, including RAG architectures, agentic systems…
…Architect end-to-end generative AI solutions with a focus on LLMs training , deployment and RAG workflows. Collaborate closely with customers to understand their…
…teams to develop scalable and generalizable manipulation solutions for deployment in robotic systems. Work with inference, deployment, and application teams to integrate deep learning…
…to drive the direction of architecture choice and deployment system design. Formulate specialized optimization solutions for various inference paradigms and scenarios (autoregressive models, denoising…
…non-engineering teams to understand their needs and build the right solutions - Architect and build services to handle write-heavy consumer traffic and data…