Jobs
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Indexed directly from employers. Every age is their own publish date.
Searching titles and descriptions for “ML Engineer, Inference & Optimization”. A match may be a passing mention rather than the job itself. Titles only.
651 roles across 745 listings · show every listing · page 1 of 27
…5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale. - Inference Mastery: Proven expertise in inference optimization…
…role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and…
…analyze and optimize infrastructure efficiency for AI workloads (compute, storage, networking) MINIMUM QUALIFICATIONS: - Bachelor’s degree in Computer Science, Computer Engineering, relevant technical field…
…beyond basic A/B testing (causal inference, cohort analysis, time-series analysis) - Experience working in voice AI, ML/AI platforms, or developer tools Notice…
…Deployment, distributed inference, and optimization of LLMs - Prompt engineering and context management - FM evaluation and benchmarking - Experience with AWS AI/ML ecosystem (Amazon Bedrock…
…gains into measurable business impact Collaborate closely with engineers on model serving infrastructure (SageMaker, GPU inference, real-time feature stores) to deploy models efficiently…
…You'll partner closely with scientists developing and fine-tuning large language models, engineers building low-latency inference infrastructure, and product teams defining customer…
…researchers and engineers across the full ML lifecycle, from architecture exploration and large-scale training to post-training optimization and inference acceleration. This is…
…Apply ML/DL to semiconductor manufacturing: defect detection, inspection and metrology, yield optimization, and process control. Analyze EDA and manufacturing application architectures and find…
…PhD in Computer Science, Electrical Engineering, or related field Experience in LLM efficiency research such as efficient attention, inference acceleration, or KV cache compression…
…consumption Partner closely with AI/ML engineers and platform teams to deliver high-quality data for model training, inference, and agent workflows Implement data…
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer…
…designing and tracking ML experiments using tools such as MLflow Familiarity with edge deployment or model optimization techniques for inference (e.g., quantization, TensorRT…
…data/ML infrastructure products - Experience with WebSockets, real-time data, or custom visualization tooling - Contributions to open-source projects or public engineering work BENEFITS…
…Wayve allocates and orchestrates training and inference workloads across thousands of GPUs and multiple data centers, ensuring optimal throughput, resiliency, and cost efficiency. Petabyte…
…grade ML workloads on our unified platform, encompassing the entire MLOps lifecycle from end-to-end pipeline creation and optimization (training/inference) to seamless…
…PREFERRED SKILLS - Familiarity with developer ecosystems, APIs, and modern software development practices. - Knowledge of AI/ML infrastructure, model serving, and performance optimization concepts. BENEFITS…
…engagement - Experience utilizing AI/ML tools/techniques to develop large scale optimization models - • Lead discovery workshops with customer CTO / engineering teams to map high…
…Our work spans ML and Data science across predictive modeling, reinforcement learning (Bandits), adaptive experimentation, causal inference, data engineering. Key job responsibilities Search Supply…
…ML engineers / systems-oriented engineers working on model optimization and ML efficiency. Set Technical Direction: Define the roadmap for training optimization, inference optimization, launch…
…Discovery Mode ML squad, driving evaluation and continuous improvement of the models that power measurement and campaign optimization Partner with ML engineers to develop…
…ML systems in production environments. - Expertise in designing, building, and maintaining production-quality ML systems and infrastructure. - Experience training, serving, debugging, and optimizing large…
…LLMs, deployment and distributed inference of LLMs, RAG, FM evaluation, Vector DBs, Agentic workflows, prompt/context engineering, and MLOps. - Hands-on experience with AWS…
…A day in the life You work alongside customer engineering teams during live AI implementation sprints — debugging inference pipelines, optimizing RAG architectures, tuning agent…
…Build and enhance C++ & python backend implementations for ASR, TTS, and S2S pipelines, leveraging CUDA for GPU acceleration Optimize Inference Performance: Improve streaming latency…