Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Performance Engineer (Inference, Training & GPU)”. A match may be a passing mention rather than the job itself. Titles only.
40 roles · page 1 of 2
…custom machine learning accelerators, Inferentia and Trainium. The Acceleration Kernel Library team is at the forefront of maximizing performance for Amazon's custom ML…
…performance GPU kernels for our novel model architectures - Integrate kernels into PyTorch pipelines (custom ops, extensions, dispatch, benchmarking) - Profile and optimize training and inference…
…Background in GPU or Deep Learning ASIC architecture evaluation for training and/or inference. Strong programming skills in Python and C++. Ways to stand…
…improvement NICE TO HAVE - Experience with AI/ML inference or training infrastructure - Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…Scale & Performance: Experience training models across distributed systems (multi-GPU/multi-node) and optimising training and inference performance (e.g., XLA, Triton, CUDA, Pallas…
…Proven experience in GPU cluster scale continuous profiling & analysis tools/platforms Solid experience in large AI job performance analysis for training/inference workload Knowledge…
…Demonstrated success in profiling and optimizing performance bottlenecks across the LLM training or inference stack. Familiarity with data center-scale orchestration, cluster schedulers, or…
…Provide front-line leadership of engineering efforts to improve model performance and scale our inference and training systems Become familiar with the team’s…
…Provide front-line leadership of engineering efforts to improve model performance and scale our inference and training systems Become familiar with the team’s…
…Kernel Engineer, you'll be responsible for identifying and addressing performance issues across many different ML systems, including research, training, and inference. A significant…
…building and scaling inference systems for LLMs or multimodal models. - Have worked with GPU-based ML workloads and understand the performance dynamics of large…
…training, inferencing, KV cache, RAG and more. Utilizing the best from NVIDIA's network products, DPUs and NICs, working closely with hardware architects, SW…
…We scale the DNN models and training/inference frameworks to systems with hundreds of thousands of nodes. Optimizing communication performance: Identify and eliminate bottlenecks…
…Strong candidates should have familiarity with elements of language model training, evaluation, and inference and eagerness to quickly dive and get up to speed…
…on distributed training, real-time inference, and communication optimization across large-scale systems. Join our world-class team of researchers and engineers building next…
…KEY RESPONSIBILITIES: - Optimize system and GPU performance for high-throughput AI workloads across training and inference - Analyze and improve latency, throughput, memory usage, and…
…largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love…
…This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This…
…leverage our large-scale GPU training and inference fleet through an observable, reliable and high-performance distributed AI/GPU communication stack. Currently, one of…
…Direct experience in developing or deploying large scale GPU based AI applications, like Large Language Model, for training and inference Strong background in building…
…experience in an engineering leadership role. In-depth expertise in linear algebra and performance optimization of Deep Learning training and inference Background in parallel…
…8+ years of hands-on validated ML/DL performance engineering experience with focus on improving GPU compute efficiency of training and inferencing workloads. Experience…
…Performance Optimization Optimize inference pipelines using CUDA, TensorRT, and mixed-precision techniques for real-time performance. Implement multithreaded scheduling and containerized deployments for automotive…
…Partner with Solutions Architecture, Research, and Engineering to design solutions/POCs and prove value across inference and post-training engagements. Inform product needs by…