New See exactly what you're overpaying AWS in under 60 seconds. Try the Calculator for free

GCP cost optimization: CUDs, SUDs, and real savings

A practitioner’s guide to Google Cloud’s discount programs, which lever to pull first, and how to commit without locking in waste.
Updated September 3, 2026
20 min read
GCP cost optimization: CUDs, SUDs, and real savings
In this article
Key takeaways
1
SUDs apply automatically, with up to 30% off eligible usage.
2
Resource-based CUDs offer deep discounts for stable Compute Engine resources.
3
Compute Flexible CUD rates vary by service and SKU; eligible GKE usage receives 28% for one year and 46% for three years.
4
CUDs and SUDs do not stack on the same usage.
5
GKE needs a mode-specific strategy for Standard and Autopilot workloads.
5
Spot VMs can provide up to 91% discounts but may be preempted at any time.
5
GCP commitments cannot be cancelled, so commit to your durable baseline, not peak usage.
Google Cloud costs can quickly grow when workloads run continuously without the right pricing strategy. CUDs, SUDs, and Spot VMs can reduce compute costs, but each works best for different usage patterns.

The key is knowing which discount fits your workload and how much usage you can safely commit to. This guide explains the main GCP cost-saving options, current discount rates, and how to build a strategy around your baseline usage.

How GCP discounts work

Google Cloud offers several discount mechanisms, and each works differently. The important distinction is whether the discount is automatic, tied to specific resources, or based on eligible spend.
Discount Commitment Typical discount Best for
On-demand None 0% Spiky, unpredictable workloads
Sustained use (SUD) None, automatic Up to 30% for eligible usage Steady eligible Compute Engine workloads
Resource-based CUD 1 or 3 years Up to 55% for many eligible machine types; higher rates apply to some memory-optimized types Stable resources in a defined region
Compute Flexible CUD 1 or 3 years Varies by service and SKU; GKE: 28% for 1 year and 46% for 3 years Predictable spend across eligible compute services
Spot VMs None Up to 91% for many eligible resources Fault-tolerant and interruptible workloads
Google Cloud’s committed use discounts documentation explains that resource-based CUDs are designed around a minimum amount of specific Compute Engine resource usage in a particular region. Compute Flexible CUDs instead use a minimum hourly spend across eligible services and resources, including Compute Engine, GKE, and Cloud Run.

Sustained use discounts are automatic and can reduce eligible Compute Engine costs as usage continues during the billing month. Eligibility and maximum discount levels depend on the machine family. See Google Cloud’s sustained use discounts documentation for the applicable rules.

Most importantly, CUDs and SUDs should not be treated as additive discounts on the same usage. SUDs do not apply to resource usage already covered by CUDs. Google’s sustained use discount guidance explains how SUD eligibility interacts with other pricing mechanisms.

Choose your CUD type

The right CUD depends on how stable the underlying workload is and how much flexibility you need.

Stable Compute Engine workloads

Use a resource-based CUD when you have a durable Compute Engine baseline and the machine configuration and region are unlikely to change.

This approach can provide deeper discounts, but the commitment is tied more closely to the resources you specify. It is therefore better suited to workloads with predictable resource requirements.

For the exact resource requirements and eligible configurations, refer to Google Cloud’s committed use discounts documentation.

Variable Compute Engine workloads

Consider a Compute Flexible CUD when your baseline is predictable but the underlying resources may change.

Flexible commitments can cover eligible spend across Compute Engine, GKE, and Cloud Run. A single flexible commitment can therefore provide more room to move workloads between supported compute services and resources than a resource-based commitment.

GKE Standard

For GKE Standard, the commitment strategy depends on how the cluster consumes compute.

Standard node pools use underlying Compute Engine resources. Applicable Compute Engine resource-based CUDs can therefore cover eligible node compute. Eligible GKE Standard usage can also be covered through Compute Flexible CUDs.

See Google Cloud’s GKE committed use discounts documentation for current GKE eligibility and commitment details.

GKE Autopilot

Autopilot uses a different billing model. Autopilot CUDs are no longer available for new purchases. Google Cloud directs customers toward Compute Flexible CUDs for eligible Autopilot usage.

This makes Compute Flexible CUDs the relevant commitment model when predictable Autopilot compute spend is part of the baseline.

The Google Cloud GKE CUD documentation provides the current guidance for Autopilot commitment coverage.

Cloud Run

Cloud Run also fits into the Compute Flexible CUD model for eligible usage. However, discount rates vary by billing model and SKU, so teams should verify the applicable Cloud Run pricing rather than assuming the GKE rate applies.

See Google Cloud’s Cloud Run committed use discounts documentation for current eligibility and pricing details.

Teams running a mix of Compute Engine, GKE, and Cloud Run workloads may therefore be able to use a unified flexible commitment rather than managing completely separate compute commitments.

Spot VMs

Spot VMs are a non-commitment option for workloads that can tolerate interruption. They offer variable pricing and discounts of up to 91% off on-demand prices for many machine types, GPUs, TPUs, and Local SSDs, but Google Cloud can preempt them at any time when it needs capacity.

Google Cloud’s Spot VMs documentation explains the pricing model, preemption behavior, and supported resources.

Use Spot VMs for fault-tolerant work such as batch processing, CI runners, distributed processing, and workloads that can checkpoint, retry, or be recreated. Do not rely on them as the sole capacity for workloads that cannot withstand interruption.

Spot prices can change, and availability and preemption rates can vary by machine type and zone. Before deploying Spot capacity, check the applicable current price, availability, and historical preemption data for the machine type and location you plan to use.

Cloud SQL and BigQuery

Cloud SQL and BigQuery have their own service-specific commitment programs. They should therefore be evaluated separately from Compute Engine resource-based CUDs and Compute Flexible CUDs.

The decision should start with the actual service generating the spend, the stability of that spend, and the commitment model available for that service.

Which savings lever to use first

Start with the least risky optimization and move toward longer commitments only after the baseline is understood.
Your situation Best first lever Why
Oversized or idle VMs Rightsize first Prevents committing to waste
Eligible steady Compute Engine usage SUD Automatic savings without a commitment
Stable Compute Engine baseline Resource-based CUD Deeper discount for predictable resources
Variable but predictable compute spend Compute Flexible CUD More flexibility across eligible compute
GKE Standard node baseline Resource-based or Compute Flexible CUD Depends on resource stability and coverage needs
GKE Autopilot baseline Compute Flexible CUD Current model for eligible Autopilot commitment spend
Heavy BigQuery usage BigQuery CUD Service-specific commitment
Predictable Cloud SQL fleet Cloud SQL CUD Service-specific commitment
Fault-tolerant batch jobs Spot VMs Lower-cost interruptible capacity
The basic rule is simple: rightsize before you commit.

A commitment sized to an over-provisioned baseline does not make the waste disappear. It locks that baseline into a discounted rate.

Google Cloud’s CUD recommendations can help identify spending and usage patterns and provide recommendations for commitment levels.

Where your GCP spend concentrates

Most cloud bills cluster around a few major services, and each has different optimization mechanics.

Compute Engine

Compute Engine is often a major infrastructure cost for GCP workloads.

Resource-based CUDs are useful when the baseline is stable enough to commit to specific resources and a region. Compute Flexible CUDs provide a broader model for eligible compute spending when workloads may move across supported services or configurations.

Do not automatically commit the entire Compute Engine bill. Separate:
  • Stable baseline usage
  • Seasonal or growth-related usage
  • Bursty usage
  • Idle or oversized resources
  • Workloads that may migrate or change machine family
Commit the durable portion and leave uncertain demand outside the commitment.

Cloud SQL

Cloud SQL does not use Compute Engine SUDs. It has its own committed use discount model for eligible Cloud SQL usage. When evaluating Cloud SQL commitments, account for the actual production architecture rather than looking only at the primary instance. High-availability configurations can materially change the amount of CPU and memory capacity being consumed.

See our Cloud SQL pricing guide for a deeper look at how High Availability changes the cost calculation.

BigQuery

BigQuery has its own spend-based CUD program. It can make sense when analytics consumption is predictable enough to support a commitment.

However, do not treat every BigQuery workload as a commitment candidate. Separate recurring baseline usage from highly variable projects, experiments, and temporary analytics spikes before determining the committed amount.

See our BigQuery committed use discounts guide for a deeper look at sizing.

GKE

GKE needs special attention because Standard and Autopilot use different compute models.

GKE Standard node compute can use applicable Compute Engine resource-based CUDs, while eligible GKE Standard and Autopilot usage can use Compute Flexible CUDs.

This means teams should identify the GKE operating mode before selecting a commitment strategy.

Autopilot CUDs are no longer available for new purchases; eligible Autopilot usage can instead use Compute Flexible CUDs.

For current eligibility rules, see Google Cloud’s GKE committed use discounts documentation.

Cloud Storage

Cloud Storage does not use Compute CUDs. Storage cost optimization therefore relies on storage class selection, lifecycle management, retention policies, and reducing unnecessary stored data.

Do not evaluate a storage-class migration on storage price alone. Retrieval, operations, and network costs can change the overall economics.
Watch out:
egress is billed separately and is not covered by compute commitments. Model network costs before moving data or changing architectures solely for storage savings.

An illustrative CUD savings example

Illustrative only. Your rates will vary by machine family, region, commitment type, and coverage.

Monthly Compute Engine spend at on-demand rates
=
$50,000
Stable baseline covered by a 3-year CUD
=
80% = $40,000
Illustrative maximum discount
=
55%
Maximum gross saving vs. undiscounted rates
=
$40,000 × 55% = $22,000
Remaining 20% stays on-demand for bursts
=
$10,000
The $22,000 figure is a maximum gross comparison against undiscounted on-demand rates. It should not automatically be treated as the incremental reduction on the actual bill.

If the same usage would otherwise qualify for a sustained use discount, the incremental savings from replacing that SUD with a CUD will be lower because the discounts do not stack on the same usage.

For example, if eligible usage would otherwise receive a 30% SUD, the correct comparison is between the CUD-discounted price and the already-SUD-discounted baseline—not between the CUD and the full undiscounted price.

That distinction matters when evaluating whether a commitment is actually worth taking.

The remaining 20% deliberately stays outside the commitment so a burst or temporary growth does not force the team to pay for committed capacity it may not use.

Common commitment mistakes

Committing before rightsizing

A discounted VM is still expensive if the workload does not need the capacity.

Rightsize first, then calculate the baseline you are actually willing to commit to.

Assuming CUDs and SUDs stack

They do not provide two independent discounts on the same covered usage.

When comparing a CUD against the status quo, account for any SUD that would otherwise apply. See Google Cloud’s sustained use discount guidance for the current rules.

Treating peak usage as the baseline

Peak usage is not the same as committed usage.

If demand regularly moves between $30,000 and $50,000 per month, committing the full $50,000 simply because it is the highest observed point increases underutilization risk.

Ignoring GKE mode

GKE Standard and Autopilot do not have identical commitment mechanics.

Standard node compute can use applicable Compute Engine resource-based CUDs, while eligible GKE Standard and Autopilot usage can use Compute Flexible CUDs.

Committing during a migration

A migration can make historical usage a poor predictor of future usage.

If machine families, regions, architectures, or services are about to change, wait until the new baseline is sufficiently understood before making a long-term commitment.

How we handle commitments at Usage.ai

Native GCP commitments cannot be cancelled after purchase, which makes accurate baseline analysis important before committing. Google Cloud states that after purchasing a CUD, you cannot cancel the commitment and remain responsible for the commitment fee throughout the applicable term.

With Flex Insured Commitments, teams can capture commitment-based savings while reducing the risk of unused capacity. Our Flex Commitments reduce cloud spend by 30–50% on average across covered workloads, while qualifying commitments include cashback protection for eligible underutilization.

For GCP teams, we analyze billing data to identify stable usage that may be suitable for commitments, help determine an appropriate commitment level, and support eligible commitment actions. 

This helps teams capture more of the savings available through GCP commitments without taking on the traditional risk of unused capacity. 
Stop overpaying on GCP
See your GCP savings before you commit

Automate Compute Engine CUDs and flexible commitments with Usage.ai and get cashback protection without the risk of long-term commitments.

Frequently asked questions

Does DynamoDB Streams consume my table's read capacity?

No. Streams reads are billed in their own Streams read request units, which are entirely separate from the read capacity units your application uses against the table. Enabling and reading a stream does not consume or reduce your table's provisioned or on-demand read capacity.

Do Lambda functions always read Streams for free?

Almost always. Standard serverless Lambda functions triggered off a stream read for free, and Global Tables replication reads are free too. The one exception is a function running on Lambda Managed Instances, which is billed at the standard $0.02 per 100,000 rate. You still pay separately for Lambda requests and duration in every case.

Do TTL deletions generate stream records?

Yes. When TTL expires an item, DynamoDB writes a delete record to the stream if Streams is enabled. It carries a userIdentity field marking it as a service-initiated delete. Streams does not charge per record, but for a custom consumer a high TTL expiry rate can still raise cost by increasing GetRecords calls. You can filter to invoke your function only on those records, or drop all deletes by filtering event type.

How long are Streams records retained, and can I extend it?

Records are retained for 24 hours and then deleted automatically. There is no setting to extend that window. If a consumer falls more than 24 hours behind, it loses records. When you need a longer replay window, route changes to Kinesis Data Streams, which supports retention up to 365 days.

How do I monitor my Streams read volume for cost?

Use the CloudWatch SuccessfulRequestLatency metric with the SampleCount statistic, which reports the number of GetRecords calls per stream. Because it counts free reads (Lambda, Global Tables) alongside billable ones, subtract those to isolate charged reads, then alarm on the result. DynamoDB Streams does not support cost allocation tags, so CloudWatch is the primary native visibility path.

Share
Facebook
X
LinkedIn
Reddit
Cut cloud cost with automation
Latest from our blogs