Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Software Engineer, Inference - Performance Optimization”. A match may be a passing mention rather than the job itself. Titles only.
63 roles across 64 listings · show every listing · page 1 of 3
…AWS Neuron is the software of Trainium and Inferentia, the AWS Machine Learning chips. Inferentia delivers best-in-class ML inference performance at the…
…Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions, ensuring every FLOP counts in delivering optimal performance for…
…This role focuses on developing and optimizing compute architectures that deliver exceptional performance and efficiency for inference workloads. You will work on cutting-edge…
…You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems…
…a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products. As generative media reshapes…
…This includes everything from designing microservices, optimizing our routing algorithms, understanding road network graphs, building monitoring and analytics infrastructure, optimizing our deployment pipeline, and…
…model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning. - Debug performance and correctness issues spanning model code, compiler IRs, runtime behavior…
…the-baseten-inference-stack/ - Driving model performance optimization https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/ RESPONSIBILITIES Core Engineering Responsibilities - Design…
…including optimizing ML inference and other GPU workloads on embedded compute (e.g., CUDA, TensorRT) - Experience with Rust in production or performance-critical systems…
…THE ROLE A backend engineer at Hebbia blends expertise in systems, application layer software, and data modeling to build highly efficient software solutions. You…
…data as needed) Evaluation & Inference: Implement algorithms and software to analyse and evaluate the performance of AI models. Optimising performance of AI/ML models…
…scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads…
…functional teams — from embedded and hardware engineers to UI and cloud teams Commitment to measuring and optimizing performance; skilled at using profiling and instrumentation…
…and high-performance backend infrastructure to support distributed training, inference, and data processing pipelines. - Lead technical design discussions, mentor other engineers, and establish best…
…multi-node inference, intra-node execution, state management, and robust error handling. - Optimize routing and communication layers using our collectives. - Utilize performance profiling and…
…high-performance data center products. As a PCB Layout Engineer at Etched, you will play a crucial role in designing and optimizing printed circuit…
…Experience in optimizing model performance through efficient processing, visualization, and analysis of large datasets. Strong software development skills and Python/C++ coding proficiency. Ability…
…Experience working on LLM inference pipelines, transformer model optimization, or model-parallel deployments. Demonstrated success in profiling and optimizing performance bottlenecks across the LLM…
…Partner with infrastructure engineers to develop and optimize systems for training, inference, monitoring, and deployment. Explore new ideas at the edge of what’s…
…About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part…
…This role is for a software engineer in the Distributed Training team for AWS Neuron. This role is responsible for development, enablement and performance…
We are looking for a Storage Software Architect to join the architecture group. You will be part of a team that shapes the next…
…We scale the DNN models and training/inference frameworks to systems with hundreds of thousands of nodes. Optimizing communication performance: Identify and eliminate bottlenecks…
…time inference, and communication optimization across large-scale systems. Join our world-class team of researchers and engineers building next-generation software and hardware…
…Integrate and enable new machine learning models into the existing platform or client environments. - Performance Optimizations: Improve system performance, efficiency, and scalability of deployed…