Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “Compiler Engineer - AI Inference”. A match may be a passing mention rather than the job itself. Titles only.
296 roles across 342 listings · show every listing · page 5 of 12
…GitHub) Experience with PyTorch, TensorFlow or similar machine learning toolsets Experience or knowledge of training/inference of Large scale AI models - CV and/or…
…world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10…
…The emphasis is on AI Data Center infrastructure software, accelerator enablement, workload orchestration, inferencing optimization, workload deployment, platform readiness, and regional product positioning for…
…Experience in machine learning (ML) infrastructure development or ML performance engineering. Preferred qualifications: Experience with ML compilers and their internals, experience writing compiler optimization…
…The profiler plays a crucial role to internal and external customers in optimizing AI workloads across hardware platforms such as Trainium and Inferentia devices…
…GPU performance work — CUDA/Triton kernels, torch.compile, operator fusion, quantization — and interest in inference-efficiency domains such as AI-RAN. Experience benchmarking AI…
…model training and inference. Solid software engineering fundamentals and system architecture thinking, with the ability to build modules and drive engineering practices in complex…
…Deep Learning Training and Inference Frameworks, with the goal of supporting NVIDIA's top AI researchers and software engineers in driving the future of…
…AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices, automotive, and robotics. The compiler must…
…engineer in the areas of Linux User-space for Machine Learning Use cases. The development target is Qualcomm high-performance inference accelerator AI100 and…
…Model compilation and systems optimization Model compression techniques (quantization, distillation) Efficient model architecture design Scalable inference systems Collaborating cross-functionally with research, engineering, and…
…an active compiler toolchain codebase, such as LLVM, MLIR, GCC, MSVC, Glow Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration…
…Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization. Partner with systems engineers to integrate models into Cloudflare…
…engineering, performance engineering, ML systems engineering, infrastructure engineering, or related areas. - Hands-on experience with large-scale GPU or accelerated computing infrastructure for AI…
…cloud, and with workload orchestration across them. - Experience with inference optimization techniques (quantization, distillation, compilation, or runtime tuning) for production serving. - A track record…
…inference systems, compiler/runtime tooling, and hardware-aware optimization techniques. YOU MAY BE A FIT IF YOU HAVE - Strong systems engineering experience in AI…
…inference systems, compiler/runtime tooling, and hardware-aware optimization techniques. YOU MAY BE A FIT IF YOU HAVE - Strong systems engineering experience in AI…
…software) for inference or training solutions. • Develops optimized software to enable AI models deployed on hardware (e.g., machine learning kernels, compiler tools, or…
…of experience in software engineering with deep specialization in one or more ML systems domains including AI infrastructure, ML compilers, high-performance computing, GPU…
…software engineering with a focus on machine learning systems, AI infrastructure, or high-performance computing Experience developing and optimizing ML training or inference pipelines…
…As an Embedded AI Engineer, you will take Deepgram's models and make them run — fast, accurately, and efficiently — on resource-constrained embedded and…
…firmware/software co-design, DSP programming, hardware accelerator integration, or silicon validation Experience with on-device ML inference, model optimization (quantization, compilation for custom…
…silicon engineering, hardware design and verification, software, and operations. We've delivered AWS Nitro, ENA, EFA, Graviton, F1 EC2 Instances, AWS Neuron, Inferentia and…
…You will work closely with major Chinese CSPs to address their critical demands on large-scale AI training/inference, Agentic AI, gaming AI, and…
…We build automated, data-driven workflows to detect, explain, and prevent performance regressions across key deep learning workloads, partnering closely with kernel developers, compiler…