Briefing

AWS Building Blocks for Foundation Model Training and Inference

ai-dev

Use P6‑b300.48xlarge with EFA v4 and FSx for Lustre to achieve the highest collective performance and durable checkpoints for large‑scale training.

What to do now

Migrate training workloads to P6‑b300.48xlarge with EFA v4 and FSx for Lustre, then benchmark collective ops to confirm the expected 18 % speedup.

Summary

Amazon has detailed the building blocks that underpin large‑scale foundation‑model training and inference on AWS. The post maps the three core infrastructure layers—accelerated compute, high‑bandwidth networking, and distributed storage—onto the AWS ecosystem, highlighting the newest P5 and P6 EC2 families. P5 instances ship with H100 GPUs (p5.48xlarge has eight H100s, each with 80 GB HBM3 and 3.35 TB/s bandwidth) while P6 introduces Blackwell B200 and B300 GPUs that deliver 2.25 PFLOPS dense BF16 throughput and 1.44–2.10 TB of HBM3e per node. Networking is split into intra‑node NVLink/NVSwitch and inter‑node Elastic Fabric Adapter (EFA), with EFA v4 on P6 adding an 18 % collective‑performance boost over v3. Storage is layered from local NVMe SSD (30.72 TB raw per instance) to Amazon FSx for Lustre, which offers terabytes‑per-second throughput and sub‑millisecond latency, and finally to durable S3 checkpoints via Data Repository Associations.

The article also explains how these components fit into the broader OSS stack—Slurm or Kubernetes for orchestration, PyTorch or JAX for training, and Prometheus/Grafana for observability—so that engineers can align their tooling with AWS’s hardware capabilities. By presenting concrete performance numbers and deployment patterns, the post equips ML teams to make informed decisions about instance choice, network configuration, and storage hierarchy for both pre‑training and inference workloads.

Key changes

  • P5‑p5.48xlarge offers eight H100 GPUs, each with 80 GB HBM3 and 3.35 TB/s bandwidth
  • P6‑b300.48xlarge introduces Blackwell B300 GPUs delivering 2.25 PFLOPS dense BF16 throughput and 2.10 TB HBM3e per node
  • EFA v4 on P6 provides 800 GB/s aggregate bandwidth, an 18 % improvement over EFA v3
  • Local NVMe SSD per instance delivers 30.72 TB raw capacity for hot data
  • FSx for Lustre delivers terabytes‑per‑second throughput, millions IOPS, and sub‑millisecond latency, with S3 integration for checkpoint persistence
  • EFA v3 on P5en reduces packet latency by ~35 % compared to EFA v2

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting