AWS Inferentia and Annapurna Labs
AWS Inferentia is an AWS chip product line. Annapurna Labs, Amazon's semiconductor division acquired in 2015, builds Nitro, Graviton, and Trainium and ranks among TSMC's top five fabless customers. Do not assume Annapurna designs every AWS accelerator.
WHY IT EXISTS Running deep learning inference at cloud scale on general purpose GPUs is expensive, and GPU supply is a bottleneck shared across every cloud customer competing for the same chips. Amazon addressed both problems by acquiring Annapurna Labs in January 2015 and turning it into an in house chip design arm. Owning silicon design lets AWS build accelerators tuned specifically for its own instance types and pricing model, instead of depending entirely on third party GPU vendors and their roadmaps.
THE MENTAL MODEL Picture Annapurna Labs as Amazon's private chip design studio, separate from the foundries that physically manufacture the silicon. Annapurna draws the blueprints, TSMC and other fabs cut the wafers. That fabless model is the same one Apple, Qualcomm, and Nvidia use, design in house, manufacture out of house, and Annapurna operates at large enough volume to rank among TSMC's top five customers, giving Amazon negotiating leverage and manufacturing priority a smaller in house team could not command.
HOW IT WORKS Annapurna Labs was founded in 2011 by Hrvoje Bilic and Nafea Bshara, then acquired by Amazon in 2015 and folded in as a wholly owned subsidiary. Its confirmed product lines include the Nitro system, which offloads virtualization, networking, and storage from the host CPU onto dedicated hardware, Graviton, Amazon's Arm based general purpose server processors, and Trainium, built for training large models. AWS Inferentia sits alongside these as AWS's purpose built inference chip line, accessed through EC2 Inf1 and Inf2 instances, designed to run already trained models cheaply and efficiently rather than to train them.
WHEN IT MATTERS This matters when you are choosing EC2 instance types for a production inference workload with predictable, high volume traffic, translation, recommendation, image classification, where Inferentia backed instances typically undercut GPU instances on cost per inference. The footgun is assuming every custom AWS chip shares one architecture or one toolchain. Nitro, Graviton, Trainium, and Inferentia solve different problems, so compatibility and performance claims about one should never be assumed to carry over to another without checking the specifics.
ONE CONCRETE EXAMPLE A team running a recommendation model on GPU backed EC2 instances migrates inference to Inf2 instances, compiling their PyTorch model with the AWS Neuron SDK. They see lower cost per thousand inferences and comparable latency, but training stays on GPU or Trainium backed instances instead, since Inferentia is scoped specifically to inference, the same kind of specialization that separates each custom chip line AWS has built out.
Read the original → en.wikipedia.org
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.