Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Finance Manager, GPU, Workload Health”. A match may be a passing mention rather than the job itself. Titles only.
43 roles across 45 listings · show every listing · page 1 of 2
…Development experience with HPC frameworks for workloads in a domain such as oil and gas, automotive / aerospace, financial services or pharmaceuticals. Proficient in one…
…years of management of technical, enterprise customer facing resources or equivalent experience - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU…
…GPU delivery, health monitoring, triage automation, and diagnostic services. These are essential for running distributed AI/ML/HPC workloads across thousands of GPUs, leveraging…
…The Managed Inference platform is where customers run production LLM workloads without managing low-level infrastructure, and it is one of Crusoe's fastest…
…caching, or GPU-backed workloads. - Experience building internal tooling for workflow inspection, replay, evaluation, or model debugging. - Experience defining SLOs, managing error budgets, designing…
…These high-performance fabrics are the foundation for frontier AI customers and tier-0 workloads running on Oracle Cloud Infrastructure. As a Senior Manager…
…You'll own the products that turn raw GPU hosts into reliable, production-ready AI compute - driver and firmware management, fleet-wide health and…
…Preferred Qualifications Experience optimizing large-scale GPU inference or training workloads for latency, throughput, utilization, availability, and cost. Experience building or operating model serving…
…Infrastructure & AI COGS Ownership Own financial visibility into compute, bandwidth, storage, and AI-related costs (e.g., inference, GPU/accelerated workloads) Translate product and…
…You Have 7+ years of product management experience building resource orchestration, workload management platforms or large-scale distributed systems Track record of building developer…
…training and inference workloads. Expertise in GPU networking technologies including GPUDirect RDMA and GPU-aware communication stacks. Experience with congestion management, adaptive routing, traffic…
…AI workloads. Develop congestion management, load balancing, resiliency, and failover capabilities for RDMA-based networks. Analyze and improve communication performance across networking, GPU, and…
…Design, develop, and maintain production software supporting OCI network deployment and lifecycle management. Build and evolve engineering platforms, APIs, automation services, and operational systems…
…Proven ability to drive alignment across sales, operations, finance, and legal in a deal structuring capacity, managing complex internal stakeholder environments without losing momentum…
…managers, banks, or other enterprise finance organizations, with an understanding of their unique technical, operational, and business-critical requirements. Familiarity with running AI workloads…
…AI, GPU and HPC services, and support major tier-0 vendors in the generative AI industry. If you're running an AI workload at…
…This role is for AI GPU/HPC RDMA Networks. AI2CNE strives to be a global leader in the RDMA cluster networking domain and enable…
…companies (CoreWeave, Lambda, etc.). - Familiarity with high-density AI workloads (liquid cooling, >30kW racks, GPU clusters). - Experience with: - Utility engagement and power delivery constraints…
…In this role, you'll work closely with customers to design, implement, and manage AWS AI/ML and GenAI solutions that meet their technical…
…AI infrastructure components like GPU control plane and GPU data plane that provide computing resources to customer AI workloads. Manage a team that designs…
…yaml), Preview Environments, private services, and managed Postgres - Experience with AI/ML workloads: LLM inference endpoints, GPU compute, vector databases, or model serving—a…
…This role is for AI GPU/HPC RDMA Networks. AI2NE strives to be a global leader in the RDMA cluster networking domain and enable…
…lifecycle management. • Strong knowledge of cloud infrastructure, GPU/HPC capacity, networking, storage, observability, platform operations, security, and service management for high-scale workloads. • Demonstrated…
…The Strategic Customers Engineering team manages Oracle Cloud Infrastructure’s most important and fastest-growing customer relationships. These customers are building workloads on OCI…
…Deep understanding of financial modeling for Capex-heavy unit economics. - Intellectual Curiosity: A student of the AI stack—from GPU power requirements to the…