Supported storage services for Cluster Director

This document provides an overview of the Google Cloud storage services that you can use in Cluster Director. By learning the different storage solutions that you can use in your cluster, you can select the best storage options for your workload requirements and budget.

High-performance storage is a critical component for large-scale artificial intelligence (AI) and high performance computing (HPC) workloads. Cluster Director supports two distinct storage categories:

  • Shared network storage: cluster-wide file systems and object stores that all nodes share to load training datasets, process data, save model checkpoints, and host user home directories (/home).

  • Node boot disk storage: compute instance-level block storage that powers the Linux operating system (OS), GPU drivers, and Slurm daemons on individual compute and login nodes.

Shared network storage

Shared storage services connect across the cluster network to provide a unified file system or object store for all cluster nodes.

Shared file system requirement for Slurm

When you deploy a cluster with Slurm as the orchestrator, you must provide a shared file system for the /home directory on all compute and login nodes. To meet this requirement, you can configure Cluster Director to create a new Filestore or Google Cloud Managed Lustre instance, or you can use an existing instance.

Supported shared storage services

In addition to a mandatory /home file system, you can attach other shared storage solutions to your cluster. Cluster Director supports the following shared network storage services:

Storage service Features Recommended for
Filestore Overview: Filestore is a managed, high-performance NFS file storage service. This service provides a file system interface and serves as the default option for home directories for clusters in Cluster Director.
  • Home directories
  • General-purpose file storage
  • Workloads that require an NFS interface
Managed Lustre Overview: Managed Lustre is a high-performance, managed parallel file system that is optimized for AI and HPC applications. With ultra-low latency and full POSIX support, this service provides an ideal solution to migrate on-premises AI workloads to Google Cloud.

Dynamic tier support: you can create new or use existing Managed Lustre instances that use the Dynamic tier when you create or modify a cluster. If you use this tier, then your cluster must use an existing VPC network that has the dynamic_tier_capacity quota allocated.
  • Home directories
  • Migration of AI or machine learning (ML) workloads to Google Cloud
  • Model simulations
  • Workloads with frequent small reads and writes
  • High-throughput burst workloads
Cloud Storage

Overview: Cloud Storage is a scalable, durable, and cost-effective object store. When you create a cluster, you can select from the following storage classes.

  • Rapid storage class: provides low latency and high throughput for AI and ML workloads.
  • Standard storage class: balances cost and performance for general-purpose use cases.
  • Autoclass: automatically adjusts object storage classes based on workload access patterns to reduce costs.

Through integration with Cloud Storage FUSE, you can mount Cloud Storage buckets as local file systems for model checkpoints and training data.

  • High-performance data access with the Rapid storage class
  • Cost-effective data storage with Standard and other storage classes
  • Data processing and preparation
  • Model training data and checkpoints

Node boot disk storage

Each controller, login, and compute node in a cluster requires a boot disk for the operating system (OS), GPU drivers, and system daemons. Boot disks attach to individual nodes and don't share files across the cluster.

Cluster Director supports the following boot disk options:

Boot disk option Features Recommended for
Standalone boot disks Overview: Cluster Director provisions a dedicated Persistent Disk or Hyperdisk volume for each node. This configuration is the default choice when you create a cluster.
  • Standard cluster deployments
  • Static clusters with fixed node counts
  • Workloads with predictable, constant node usage
Hyperdisk storage pools1 Overview: Hyperdisk storage pools let you share capacity and performance across boot disks in a specific zone to optimize costs. In Cluster Director, you can use existing Hyperdisk pools or Hyperdisk Exapools for boot disks only. Because boot disks require Hyperdisk Balanced volumes, you can't use storage pools that use Hyperdisk Throughput volumes.
  • Boot disks for compute and login nodes
  • Clusters that scale compute nodes on demand
  • Workloads with shared capacity and IOPS requirements

1 Only machine series that support Hyperdisk Balanced volumes can use storage pools in Cluster Director. If you use storage pools with an unsupported machine series, then the cluster creation or modification operation fails. For more information, see Machine series support for Hyperdisk.

What's next