HyperPod
HyperPod runs model customization on persistent compute clusters that remain available across multiple training jobs. HyperPod provides automatic fault detection and recovery, making it ideal for large-scale and long-running training workloads.
For ephemeral training jobs, see SageMaker AI Training Jobs. For fully managed serverless training, see Serverless model customization.
For supported customization techniques and training types, see Customization techniques.
Overview
Regional availability
Available in all HyperPod supported regions.
Orchestration options
HyperPod with Amazon EKS — Interfaces: Recipes, HP-CLI. Use Kubernetes-based workload orchestration with dynamic capacity management and containerized training jobs.
HyperPod with Slurm — Interfaces: Recipes. Use Slurm-based job scheduling with head, login, and worker node configuration.
Supported techniques and training types
HyperPod supports all customization techniques and training types. See Customization techniques and Training types.
Getting started
If this is your first time using HyperPod, see Prerequisites for using SageMaker AI HyperPod to set up your cluster.
To run model customization on an existing HyperPod cluster:
Clone the recipes repository:
git clone --recursive https://github.com/aws/sagemaker-hyperpod-recipes.git cd sagemaker-hyperpod-recipes pip3 install -r requirements.txtChoose a recipe from
recipes_collection/recipes/fine-tuning/.Configure cluster-specific settings (see Cluster-specific configurations).
Launch using the appropriate method:
EKS: Use the launcher script or HP-CLI
Slurm: Use the launcher script
For detailed tutorials and the full recipe catalog, see SageMaker AI Recipes documentation.
For more information about HyperPod, see Amazon SageMaker AI HyperPod.