SOURCE-LINKED INTELLIGENCE
Fault tolerant distributed training on Amazon EKS using NVRx
Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.
Read original source ↗ Open in workspace
- recordType
- article
- region
- Global
Evidence & attribution
- AWS Artificial Intelligence Blog · 2026-09-16T18:59:25.000Z
First collected: 2026-09-19T20:26:46.936Z. This is not the publication date.