π Local Job Near You
Machine Learning Engineer, AWS Neuron Inference, Annapurna ML
Amazon
π
Seattle, United States
Location
Seattle
Posted
June 03, 2026
Commute
Local Area
Local Opportunity Near You!
This job is in your area. Enjoy a short commute and work close to home.
Job Description
Description
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine
learning accelerators and the Trn2 and future Trn3 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron.
This role develops, enables and performance tunes building blocks for all key ML model families, including Llama3, GPT OSS, Qwen3, DeepSeek and beyond.
The Neuron Inference Technology team works side by side with the Inference Model Enablement, compiler runtime engineers to create, build and tune high-performance distributed inference solutions for the latest generation Trainium accelerators. Experience optimizing LLM inference performance with kernels, Python, PyTorch or JAX is a must.
Key job responsibilities
This team develops optimized building blocks for the Neuron distributed inference library, tuning them to ensure highest performance and maximize efficiency r...
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine
learning accelerators and the Trn2 and future Trn3 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron.
This role develops, enables and performance tunes building blocks for all key ML model families, including Llama3, GPT OSS, Qwen3, DeepSeek and beyond.
The Neuron Inference Technology team works side by side with the Inference Model Enablement, compiler runtime engineers to create, build and tune high-performance distributed inference solutions for the latest generation Trainium accelerators. Experience optimizing LLM inference performance with kernels, Python, PyTorch or JAX is a must.
Key job responsibilities
This team develops optimized building blocks for the Neuron distributed inference library, tuning them to ensure highest performance and maximize efficiency r...