New See exactly what you're overpaying AWS in under 60 seconds. Try the Calculator for free →

Google Cloud GPU Pricing: NVIDIA Options, Costs, and Savings

What Google Cloud NVIDIA GPU families cost in 2026, which fits your workload, and how Spot pricing, reservations, and committed use discounts can reduce costs.
Updated October 5, 2026
15 min read
Google Cloud GPU Pricing: NVIDIA Options, Costs, and Savings
In this article
Key takeaways
1
Google Cloud offers NVIDIA GPUs from L4 for inference to H100, H200, and Blackwell systems for advanced AI workloads.
2
H100 VM pricing can reach about $88–$93/hour on demand; Spot, Flex-start, reservations, and eligible CUDs provide other consumption options.
3
GPU discount eligibility varies by family, so the right commitment depends on the GPU series, region, workload stability, and capacity needs.
Google Cloud offers NVIDIA GPUs for everything from inference and fine-tuning to large-scale model training. This guide covers the major GPU options, pricing models, and practical ways to reduce AI infrastructure costs.

It also explains when to use on-demand, Spot VMs, Flex-start, reservations, and committed use discounts so you can balance GPU performance, availability, and cost.

Short answer

Google Cloud’s accelerator-optimized portfolio includes G2 (L4), A2 (A100), A3 (H100/H200), A4 (B200), A4X/A4X Max (GB200/GB300), and G4 (RTX PRO 6000). NVIDIA T4, P4, P100, and V100 GPUs can also be attached to supported N1 machine types.

In us-central1, example on-demand VM prices range from about $0.71/hour for g2-standard-4 to about $88.49/hour for a3-highgpu-8g. Pricing varies by region and machine configuration, while newer Blackwell systems can use reservation, Spot, or Flex-start consumption models rather than standard on-demand pricing.

NVIDIA GPU options in 2026

Google Cloud’s NVIDIA portfolio spans different performance and cost levels:

G2 / L4: Designed for inference, graphics, media processing, and smaller AI workloads. The L4 has 24 GB of GPU memory.

A2 / A100: Suitable for training and fine-tuning workloads requiring more GPU memory and compute.

A3 / H100: Built for large-scale AI training and demanding inference. H100 GPUs provide 80 GB of GPU memory.

A3 Ultra / H200: Provides 141 GB of GPU memory for memory-intensive inference and training workloads.

G4 / RTX PRO 6000 Blackwell: Designed for inference, visual computing, graphics, and workloads that can benefit from fractional GPU configurations.

A4 / B200: Targets frontier AI training and large-scale workloads.

A4X / A4X Max: Uses GB200 and GB300 NVL72 systems for large AI-factory and distributed workloads.

For the latest configurations, see Google’s accelerator-optimized machine types.

Google Cloud GPU pricing

Google Cloud bills accelerator-optimized machine types at the instance level. The price includes the attached GPU plus predefined vCPU, memory, and bundled Local SSD where applicable. Rates vary by region and configuration, so check Google Cloud GPU pricing before provisioning.
Series Example GPU Example VM price Best for
G2 1× L4 ~$0.71/hr Inference, media, light GenAI
A2 1× A100 40 GB ~$3.67/hr Training, fine-tuning
A2 Ultra 1× A100 80 GB ~$5.03/hr High-memory workloads
G4 1× RTX PRO 6000 ~$4.50/hr Visual computing, inference
A3 High 8× H100 80 GB ~$88.49/hr LLM training
A3 Mega 8× H100 80 GB ~$93.40/hr Large-scale LLM workloads
A3 Ultra 8× H200 141 GB ~$84.81/hr Memory-intensive AI
A4 8× B200 Consumption varies Frontier training
These are benchmark examples, not universal rates. Google Cloud pricing varies by region, machine type, and consumption model.

Match the GPU to the workload

The cheapest GPU is not always the cheapest option overall. Choose based on model size, VRAM requirements, scaling needs, and workload duration.

Inference and smaller models: Start with G2/L4.

Training and fine-tuning: Consider A2/A100 when 40 GB or 80 GB is sufficient.

Large LLM training: Use A3/H100 when you need high memory bandwidth and distributed scale.

Memory-intensive workloads: A3 Ultra/H200 provides more GPU memory.

Visual and fractional workloads: G4 can provide a more right-sized option.

Frontier-scale training: A4 and A4X families target the largest AI workloads.

For a cross-cloud comparison, see our guide to GPU cost optimization across AWS, GCP, and Azure.

Watch out for idle GPUs

Running a GPU VM continues to generate compute charges even when the workload is idle. For development and batch environments, use autoscaling, scheduled shutdowns, and appropriate Spot capacity.

Persistent disks and other attached resources can continue generating charges after a VM is stopped. The same discipline applies to GKE GPU node pools running Spot and CUDs.

How GPU discounts work

Google Cloud provides several consumption options, but eligibility depends on the machine series.

Sustained use discounts apply automatically to eligible N1 GPU configurations. They do not generally apply to accelerator-optimized A2, A3, A4, A4X, G2, or G4 GPU instances.

Compute flexible CUDs can apply to eligible G2 and G4 usage. Other accelerator-optimized families use resource-based commitments and supported reservation options instead. Check Google’s accelerator-optimized machine documentation for current eligibility.

Resource-based CUDs can provide discounted pricing for committed resources. For accelerator-optimized GPU resources, resource-based commitments can require attached reservations. A commitment provides a pricing discount; a reservation addresses capacity availability.

Spot VMs offer significant discounts for workloads that can tolerate interruption. Google Cloud states that Spot prices can provide discounts of up to 91% for many eligible machine types and GPUs, but pricing is variable and capacity is not guaranteed. See Google Cloud Spot VM pricing for current rates.

A practical decision framework:

On-demand: Unpredictable or interruption-sensitive workloads.

Spot: Fault-tolerant training, batch processing, and checkpointed workloads.

Flex-start: Workloads that can tolerate delayed provisioning in exchange for discounted compute.

Reservations: Workloads where GPU capacity availability matters.

Resource-based CUDs: Stable usage tied to specific resources and regions.

Compute flexible CUDs: Eligible baseline spend where flexibility matters.

For the economics of longer commitments, see our guide to sustained use vs. committed use discounts.

Example: L4 cost optimization

Suppose two L4 GPUs run continuously on a g2-standard-24 instance for 730 hours. At approximately $2.00/hour, the monthly compute cost is about $1,460 before other charges.

If the workload qualifies for Spot pricing, the effective cost can be substantially lower. Google Cloud’s current pricing shows Spot rates can change and provide discounts of up to 91% for eligible resources.

For predictable G2 usage, an eligible compute flexible CUD can provide additional savings. The best option depends on workload stability and whether capacity must be guaranteed.

Avoid GPU commitment risk

GPU commitments can reduce rates, but they can also create waste when workloads change.

For example, committing to a specific A100 configuration and later moving the workload to H100 or Blackwell can leave the original commitment underutilized. Resource-based commitments are tied to specific resources and regions, so commitment terms should reflect how confident you are in your long-term GPU requirements.

With  Flex Insured Commitment Program teams can get the 30–50% savings of a 1- or 3-year commitment with none of the commitment risk.

There is no multi-year lock-in and no upfront payment: we purchase and manage the commitments on your behalf, and cashback protection covers any underutilization in real money, not credits.

Setup happens at the billing layer, with no infrastructure or code changes required, and our fee is a percentage of realized savings only: if we don’t save you anything, you pay nothing.

Learn how automated cloud commitments work.

What changed in the NVIDIA partnership?

Google Cloud’s NVIDIA portfolio has expanded from L4 and H100 systems to Blackwell-based infrastructure.

G4 provides NVIDIA RTX PRO 6000 Blackwell GPUs for workloads such as inference, graphics, and visual computing. Google Cloud has also expanded its accelerator-optimized portfolio with A4 B200 systems and A4X/A4X Max GB200 and GB300 NVL72 systems.

The result is a broader GPU selection, but also a more complicated pricing and capacity landscape. Choosing the right GPU now requires evaluating both technical fit and consumption model.
Stop overpaying on GPU compute
Turn GCP GPU commitments into savings with less lock-in.

Automate GCP GPU commitments with protection against eligible unused spend.

Frequently asked questions

Is the NVIDIA T4 still available?

Yes. T4 remains available as an attachable GPU on supported N1 machine types. The P100, not the T4, is the GPU approaching its published end-of-support date in September 2026. See Google's T4 end-of-support timeline and P100 end-of-support timeline for current dates.

Can I run a single H100 on demand?

Standard on-demand H100 configurations are centered on eight-GPU A3 High and A3 Mega shapes. Smaller A3 High configurations use other consumption options such as Spot or Flex-start rather than standard on-demand pricing.

Do sustained use discounts apply to A100 or H100?

No. Sustained use discounts apply to eligible N1 configurations. A2/A100 and A3/H100 workloads use other commitment and consumption options, including resource-based commitments and reservations where supported.

How much cheaper is Spot?

Spot VMs can provide discounts of up to 91% compared with standard on-demand pricing for many eligible machine types and GPUs. Actual discounts vary, and Spot VMs can be reclaimed when Google Cloud needs the capacity.

Vertex AI or Compute Engine for GPUs?

Compute Engine provides direct access to GPU-backed VMs that you manage. Vertex AI is a managed ML platform with its own pricing model. Compare the complete workload cost—including infrastructure, platform services, storage, networking, and operations—before choosing between them.

Share
Facebook
X
LinkedIn
Reddit
Cut cloud cost with automation
Latest from our blogs