## Amazon AWS Hits General Availability for Trn3 UltraServers: 362 FP8 PFLOPS Across 144-Chip Domains 🚀☁️⚡ Building directly on its massive custom-silicon momentum, Amazon Web Services (AWS) has officially announced the general availability of **Amazon EC2 Trn3 UltraServers**—powered by its fourth-generation, 3nm **Trainium3** ASICs. Designed specifically to tackle the punishing token-economics and memory demands of next-generation agentic workflows, multi-step reasoning models, and real-time video generation, the Trn3 platform introduces massive architectural leaps in scale-up density and fabric connectivity. --- ### 1. Under the Hood: Trainium3 and the Trn3 Gen2 UltraServer 🔬🧬 Fabricated on TSMC’s advanced 3nm process, the Trainium3 chip serves as the heavy-duty engine behind the new instances, bringing massive density upgrades over Trainium2: * **Raw Compute Power:** Each individual Trainium3 chip delivers up to **2.52 PFLOPS of FP8 compute** and natively supports high-efficiency compressed formats like **MXFP8 and MXFP4**. * **Massive HBM3e Memory:** Equipped with **144 GB of HBM3e memory per chip** (delivering 4.9 TB/s of bandwidth per device), allowing massive model weights and sprawling KV-caches to sit locally in memory. * **The Gen2 UltraServer Scale-Up Domain:** The flagship Trn3 Gen2 UltraServer integrates a staggering **144 Trainium3 chips** into a single unified scale-up domain, delivering a combined **362 FP8 PFLOPS of compute**, **20.7 TB of total HBM3e memory**, and **706 TB/s of aggregate memory bandwidth**. --- ### 2. The NeuronSwitch-v1 Fabric: Eradicating Communication Bottlenecks 🌐🔌 Scaling over a hundred high-performance chips in a single cluster creates massive interconnect hurdles. To solve this, AWS engineered a brand-new internal networking topology: * **All-to-All Connectivity:** Trn3 UltraServers replace older 2D-torus layouts with the **NeuronSwitch-v1** all-to-all fabric, doubling inter-chip interconnect bandwidth compared to previous generations. * **EC2 UltraClusters 3.0 Integration:** These individual 144-chip UltraServers can be networked together via **UltraClusters 3.0** using AWS Elastic Fabric Adapter (EFA), allowing frontier labs to scale training and inference workloads seamlessly across hundreds of thousands of custom chips without performance degradation. --- ### 3. Software and Ecosystem Readiness 🛠️🎯 To ensure smooth enterprise adoption and bypass the friction of proprietary software lock-in, AWS has deeply optimized the **AWS Neuron SDK**: * **Native PyTorch Integration:** Developers can run training and inference pipelines on Trn3 clusters without changing a single line of model code. * **Unprecedented Token Economics:** According to benchmark data on Amazon Bedrock, Trainium3 delivers up to **3x faster performance than Trainium2** while yielding over **5x higher output tokens per megawatt**, drastically reducing operating costs for high-volume enterprise inference.