This blueprint provisions a Google Kubernetes Engine (GKE) cluster with A3 High nodes (a3-highgpu-8g). A3 High VMs feature 8 NVIDIA H100 GPUs and 200 Gbps of networking throughput per GPU.
The blueprint automatically configures the following components to enable optimal GPU performance and multi-networking:
- GPU-Direct TCPX: Optimized networking stack for high-bandwidth, low-latency GPU communication.
- Multi-networking: Configures 4 secondary interfaces (VPC networks) for dedicated GPU-to-GPU traffic.
- NRI Device Injector: Automatically injects required networking and GPU configurations into your ML containers.
- Kueue and JobSet: Kubernetes-native tools for managing large-scale, multi-node training jobs with Topology Aware Scheduling (TAS).
- Cluster Toolkit: Ensure you have installed all the dependencies required in cluster toolkit and followed the setup instructions.
- Install dependencies.
- Set up Cluster Toolkit. For building the
gclusterbinary, see Install Cluster Toolkit.
- Quota: Ensure you have sufficient quota for
a3-highgpu-8gmachines in your chosen region. - IP Address: You will need the public IP address of the machine where you run
gclusterto configure the cluster's authorized networks.
Before deploying, fill out the gke-a3-highgpu-deployment.yaml file with your project-specific values:
| Variable | Description |
|---|---|
project_id |
Your Google Cloud Project ID. |
deployment_name |
A unique name for this Cluster Toolkit deployment. |
region / zone |
The GCP region and zone (e.g., us-central1, us-central1-c). |
authorized_cidr |
Your public IP address in CIDR notation (e.g., 1.2.3.4/32). |
bucket |
Name of the GCS bucket to store Terraform state. |
Option 1 (Specific Reservation) is uncommented by default in gke-a3-highgpu-deployment.yaml. To use another consumption model, comment out Option 1 and uncomment the desired option.
-
Switch to the toolkit directory:
cd ~/cluster-toolkit
-
Build the toolkit:
make
-
Deploy the infrastructure:
./gcluster deploy \ examples/gke-a3-highgpu/gke-a3-highgpu.yaml \ -d examples/gke-a3-highgpu/gke-a3-highgpu-deployment.yaml
Refer the following guide to verify GPU and networking performance using NVIDIA nccl-tests:
- Multi-Node (16+ GPUs): Multi-Node Test Plan
-
Submit the DWS Flex Start job:
kubectl apply -f examples/dws-sample-workloads/sample-job-flex.yaml
-
Consider using
kubectl get jobsandkubectl describe job <job-name>to get information about the jobs.
You can also usekubectl get podsandkubectl describe pod <pod-name>to get pod information. -
Clean up the job:
kubectl delete -f examples/dws-sample-workloads/sample-job-flex.yaml
Note: DWS Flex Start workloads require nodeSelector: cloud.google.com/gke-flex-start: "true".
-
Deploy the NCCL test JobSet:
kubectl create -f examples/gke-a3-highgpu/nccl-jobset-flex.yaml
-
Monitor pods (
kubectl get pods) and check results in the primary pod logs:kubectl logs <jobset-pod-name>
-
Clean up test resources:
kubectl delete -f examples/gke-a3-highgpu/nccl-jobset-flex.yaml
-
Submit the Queued Provisioning job:
kubectl apply -f examples/dws-sample-workloads/sample-job-flex-queue.yaml
-
Consider using
kubectl get jobsandkubectl describe job <job-name>to get information about the jobs.
You can also usekubectl get podsandkubectl describe pod <pod-name>to get pod information. -
Clean up the job:
kubectl delete -f examples/dws-sample-workloads/sample-job-flex-queue.yaml
Note: Queued Provisioning workloads require the label kueue.x-k8s.io/queue-name: dws-local-queue and annotation provreq.kueue.x-k8s.io/maxRunDurationSeconds.
-
Deploy the NCCL test JobSet:
kubectl create -f examples/gke-a3-highgpu/nccl-jobset-flex-queue.yaml
-
Monitor pods (
kubectl get pods) and check results in the primary pod logs:kubectl logs <jobset-pod-name>
-
Clean up test resources:
kubectl delete -f examples/gke-a3-highgpu/nccl-jobset-flex-queue.yaml
To avoid incurring charges for the resources created, destroy the deployment:
./gcluster destroy DEPLOYMENT_NAMENote: GCS buckets created for Terraform state are not deleted by the ./gcluster destroy command and must be deleted manually.