General GPUs overview

General GPUs are optimized for artificial intelligence (AI) and machine learning (ML) workloads that prioritize operational flexibility, resource independence, and cost-effectiveness.

General GPUs let teams deploy and scale NVIDIA accelerators with complete operational autonomy, bypassing the architectural complexity of coordinated supercomputing clusters. They are ideal for workloads that prioritize high availability, self-service provisioning, and independent scaling. Common use cases include real-time inference, RAG, prototyping, and small-to-medium model training.

Key characteristics of General GPUs

  • Operational flexibility: you can manage resources by using standard GKE (Autopilot and Standard) and Compute Engine interfaces to simplify deployment and scaling.
  • High availability: Google Cloud updates individual instances independently to help ensure updates on one node don't impact the availability of others, allowing your load balancer to redirect traffic without stalling serving fleets.
  • Networking simplicity: your workloads use standard Virtual Private Cloud (VPC) services and standard TCP/IP over Google Virtual NIC (gVNIC) interfaces to eliminate complex hardware-level fabrics.
  • Cost optimization: you can use Spot VMs for fault-tolerant research to achieve cost savings of up to 90% compared to standard on-demand rates.
  • Self-service acquisition: you can obtain capacity directly through standard reservations and on-demand requests without requiring account-team-managed workflows.

General GPU families and technical specifications

Review the workloads and hardware specifications for the General GPU machine series. For high-throughput serving, use the A3 Edge series.

A3 series (Edge and High with 1, 2, or 4 GPUs)

A3 Edge

If you are running high-throughput distributed inference workloads across multiple hosts without a coordinated cluster fabric, use the A3 Edge series to simplify deployment over standard virtual private networks.

A3 Edge machine types have NVIDIA H100 SXM GPUs and are designed specifically for serving and are available in a limited set of regions.

Attached NVIDIA H100 GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3)
a3-edgegpu-8g 208 1,872 6,000 5
  • 600: for asia-south1 and northamerica-northeast2
  • 400: for all other A3 Edge regions
8 640

A3 High (1, 2, or 4 GPUs)

If you are running inference or standard training workloads that require high performance but don't need a full 8-GPU synchronized cluster, use the A3 High machine series with 1, 2, or 4 GPUs attached.

A3 High machine types have NVIDIA H100 SXM GPUs and are well-suited for both large model inference and model fine tuning.

Attached NVIDIA H100 GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3)
a3-highgpu-1g 26 234 750 1 25 1 80
a3-highgpu-2g 52 468 1,500 1 50 2 160
a3-highgpu-4g 104 936 3,000 1 100 4 320
a3-highgpu-8g 208 1,872 6,000 5 1,000 8 640

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

G2 series

If you are deploying mainstream real-time inference applications or Retrieval-Augmented Generation (RAG) pipelines, use the G2 machine series to balance performance and cost efficiency with NVIDIA L4 GPUs.

Attached NVIDIA L4 GPUs
Machine type vCPU count1 Default instance memory (GB) Custom instance memory range (GB) Max Local SSD supported (GiB) Maximum network bandwidth (Gbps)2 GPU count GPU memory3 (GB GDDR6)
g2-standard-4 4 16 16 to 32 375 10 1 24
g2-standard-8 8 32 32 to 54 375 16 1 24
g2-standard-12 12 48 48 to 54 375 16 1 24
g2-standard-16 16 64 54 to 64 375 32 1 24
g2-standard-24 24 96 96 to 108 750 32 2 48
g2-standard-32 32 128 96 to 128 375 32 1 24
g2-standard-48 48 192 192 to 216 1,500 50 4 96
g2-standard-96 96 384 384 to 432 3,000 100 8 192

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

G4 series

If your team requires cost-effective entry-level inference or standard graphics rendering workloads, use the G4 machine series featuring NVIDIA RTX PRO 6000 GPUs.

G4 machine types use NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs (nvidia-rtx-pro-6000) and are suitable for NVIDIA Omniverse simulation workloads, graphics-intensive applications, video transcoding, and virtual desktops. G4 machine types also provide a low-cost solution for performing single host inference and model tuning compared with A series machine types.

Attached NVIDIA RTX PRO 6000 GPUs
Machine type vCPU count1 Instance memory (GB) Maximum Titanium SSD supported (GiB)2 Physical NIC count Maximum network bandwidth (Gbps)3 GPU count GPU memory4
(GB GDDR7)
g4-standard-6 6 22 0 1 20 1/8 12
g4-standard-12 12 45 375 1 20 1/4 24
g4-standard-24 24 90 750 1 20 1/2 48
g4-standard-48 48 180 1,500 1 50 1 96
g4-standard-96 96 360 3,000 1 100 2 192
g4-standard-192 192 720 6,000 1 200 4 384
g4-standard-384 384 1,440 12,000 2 400 8 768

*GPU memory is the memory available on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your accelerated workloads such as machine learning and graphics-intensive workloads.

A2 series (Standard and Ultra)

If you need to perform high-performance single-node model serving or small-scale model fine-tuning, use the A2 machine series to access dedicated NVIDIA A100 GPU resources.

A2 Ultra

Attached NVIDIA A100 80GB GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM2e)
a2-ultragpu-1g 12 170 375 24 1 80
a2-ultragpu-2g 24 340 750 32 2 160
a2-ultragpu-4g 48 680 1,500 50 4 320
a2-ultragpu-8g 96 1,360 3,000 100 8 640

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

A2 Standard

Attached NVIDIA A100 40GB GPUs
Machine type vCPU count1 Instance memory (GB) Local SSD supported Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM2)
a2-highgpu-1g 12 85 Yes 24 1 40
a2-highgpu-2g 24 170 Yes 32 2 80
a2-highgpu-4g 48 340 Yes 50 4 160
a2-highgpu-8g 96 680 Yes 100 8 320
a2-megagpu-16g 96 1,360 Yes 100 16 640

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

N1 series (NVIDIA T4 or V100)

If your priority is cost-sensitive experimentation or running low-concurrency, entry-level inference, use the N1 machine series to attach NVIDIA T4 or V100 GPUs.

NVIDIA T4

Accelerator type GPU count GPU memory1 (GB GDDR6) vCPU count Instance memory (GB) Local SSD supported
nvidia-tesla-t4 or
nvidia-tesla-t4-vws
1 16 1 to 48 1 to 312 Yes
2 32 1 to 48 1 to 312 Yes
4 64 1 to 96 1 to 624 Yes

NVIDIA V100

Accelerator type GPU count GPU memory1 (GB HBM2) vCPU count Instance memory (GB) Local SSD supported2
nvidia-tesla-v100 1 16 1 to 12 1 to 78 Yes
2 32 1 to 24 1 to 156 Yes
4 64 1 to 48 1 to 312 Yes
8 128 1 to 96 1 to 624 Yes

Capacity acquisition

General GPU accelerators use standard, self-service provisioning. You can obtain capacity through standard reservations, on-demand requests, or Spot VMs. Because on-demand capacity can experience stockouts, use Spot VMs or Flex-start where possible.

What's next