Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 

README.md

GKE H4D Blueprint

This blueprint uses GKE to provision a Kubernetes cluster and an H4D node pool, along with networks and service accounts. Information about H4D machines can be found here.

NOTE: The required GKE version for H4D support is >= 1.32.11-gke.1174000.

Create a cluster

Follow these steps to configure and deploy the GKE-H4D cluster.

Note

If you create multiple clusters using these blueprints, ensure that all VPC and subnet names are unique per project to avoid resource conflicts.

  1. Set up Cluster Toolkit. We recommend using Cloud Shell to do so because the dependencies are already pre-installed for Cluster Toolkit.

  2. Create a Cloud Storage bucket to store the state of the Terraform deployment:

    gcloud storage buckets create gs://BUCKET_NAME \
        --project=PROJECT_ID \
        --default-storage-class=STANDARD \
        --location=COMPUTE_REGION \
        --uniform-bucket-level-access
    gcloud storage buckets update gs://BUCKET_NAME --versioning
  3. In examples/gke-h4d/gke-h4d-deployment.yaml, configure the general deployment settings:

    • bucket: The name of the Cloud Storage bucket created in step 2.
    • project_id: Your Google Cloud project ID.
    • deployment_name: A unique name for your cluster deployment.
    • region: The GCP region for the cluster (e.g., asia-southeast1).
    • zone: The GCP zone for the H4D node pool (e.g., asia-southeast1-a).
    • authorized_cidr: The IP CIDR block permitted to access the Kubernetes control plane (e.g., 0.0.0.0/0 to allow all authorized users, or <YOUR-IP-ADDRESS>/32).
  4. Select a consumption model: In examples/gke-h4d/gke-h4d-deployment.yaml, select ONE consumption model from the options provided. Option 1 (Specific Reservation) is uncommented by default. To use another consumption model, uncomment the desired option and comment out Option 1.

  5. Generate Application Default Credentials (ADC) for Terraform:

    gcloud auth application-default login
  6. Deploy the blueprint:

    ./gcluster deploy examples/gke-h4d/gke-h4d.yaml -d examples/gke-h4d/gke-h4d-deployment.yaml

    When prompted, select (A)pply to provision the VPC networks, Falcon IRDMA RDMA network, service accounts, GKE cluster, and H4D node pool.


Running Workloads

DWS Flex Start Test Job

When using DWS Flex Start (Option 2 or Option 3), the node pool initializes with 0 nodes and scales up on demand when matching jobs are scheduled.

A sample batch job is provided at examples/gke-h4d/test-job-flex.yaml.

Any job applied to this node pool must meet the following requirements:

  • Flex Start Selector: Workloads must include nodeSelector: cloud.google.com/gke-flex-start: "true".

  • Tolerations: Because the h4d-pool node pool is tainted (node-type=h4d:NoSchedule) to prevent generic workloads from scheduling on HPC nodes, workloads must include the matching toleration:

    tolerations:
    - key: "node-type"
      operator: "Equal"
      value: "h4d"
      effect: "NoSchedule"

Execution and Monitoring Steps

  1. Connect to the GKE cluster:

    gcloud container clusters get-credentials <cluster-name> --region <region> --project <project-id>
  2. Submit the sample test job:

    kubectl apply -f examples/gke-h4d/test-job-flex.yaml
  3. Monitor the scale-up and execution lifecycle:

    • Check Pod Status: Initially, pods will be Pending because the H4D node pool is at size 0:

      kubectl get pods -w
      NAME              READY   STATUS    RESTARTS   AGE
      h4d-job-1-q2ksv   0/1     Pending   0          10s
      h4d-job-2-j9wla   0/1     Pending   0          10s
      
    • Inspect Autoscaler Events: Check pod events to verify GKE Cluster Autoscaler triggered provisioning for the H4D group:

      kubectl describe pods -l job-name=h4d-job-1

      Look for the TriggeredScaleUp event:

      Events:
        Type    Reason            Age   From                Message
        ----    ------            ---   ----                -------
        Normal  TriggeredScaleUp  15s   cluster-autoscaler  pod triggered scale-up by cluster-autoscaler: group h4d-pool-xxxx
      
    • Track Node Readiness: After the physical H4D VMs boot and register, the pods transition to Running:

      kubectl get nodes -w
    • Observe Completion: Once the sleep workload finishes, the pods will transition to Completed:

      kubectl get pods
      NAME              READY   STATUS      RESTARTS   AGE
      h4d-job-1-q2ksv   0/1     Completed   0          2m
      h4d-job-2-j9wla   0/1     Completed   0          2m
      
  4. Clean up the test job:

    kubectl delete -f examples/gke-h4d/test-job-flex.yaml

Note

Since the node pool is configured with max_run_duration: 900 (15 minutes), any provisioned nodes will be terminated by GKE after 15 minutes, or scaled down to 0 by Cluster Autoscaler when idle.


Run a test using the MPI Operator

The MPI Operator is installed on the cluster during the deployment. To run a test using the MPI Operator on the GKE H4D cluster, refer to https://github.com/GoogleCloudPlatform/kubernetes-engine-samples/tree/main/hpc/mpi.

Clean Up

To destroy all resources associated with the deployment, run:

./gcluster destroy CLUSTER_NAME

Replace CLUSTER_NAME with the deployment_name specified in your deployment file.

Note: GCS buckets created for Terraform state storage are not deleted by ./gcluster destroy and must be removed manually if no longer needed.