Collection of best practices, reference architectures, model training examples and utilities to train large models on AWS.
-
Updated
Aug 8, 2026 - Shell
Collection of best practices, reference architectures, model training examples and utilities to train large models on AWS.
This collection of helper scripts/ and guides for AWS SageMaker HyperPod and ParallelCluster makes it easy to get started with large-scale distributed training on Slurm-based HPC clusters and Kubernetes-based EKS clusters, as well as AI/ML model inference deployment.
Infrastructure deployment automation of SageMaker HyperPod clusters based on EKS and SLURM orchestration and Protein Language ESM-2 model training job definitions including NVIDIA BioNemo
Experimental scripts for Amazon SageMaker HyperPod
Add a description, image, and links to the hyperpod topic page so that developers can more easily learn about it.
To associate your repository with the hyperpod topic, visit your repo's landing page and select "manage topics."