The key is knowing which discount fits your workload and how much usage you can safely commit to. This guide explains the main GCP cost-saving options, current discount rates, and how to build a strategy around your baseline usage.
How GCP discounts work
Google Cloud offers several discount mechanisms, and each works differently. The important distinction is whether the discount is automatic, tied to specific resources, or based on eligible spend.| Discount | Commitment | Typical discount | Best for |
|---|---|---|---|
| On-demand | None | 0% | Spiky, unpredictable workloads |
| Sustained use (SUD) | None, automatic | Up to 30% for eligible usage | Steady eligible Compute Engine workloads |
| Resource-based CUD | 1 or 3 years | Up to 55% for many eligible machine types; higher rates apply to some memory-optimized types | Stable resources in a defined region |
| Compute Flexible CUD | 1 or 3 years | Varies by service and SKU; GKE: 28% for 1 year and 46% for 3 years | Predictable spend across eligible compute services |
| Spot VMs | None | Up to 91% for many eligible resources | Fault-tolerant and interruptible workloads |
Sustained use discounts are automatic and can reduce eligible Compute Engine costs as usage continues during the billing month. Eligibility and maximum discount levels depend on the machine family. See Google Cloud’s sustained use discounts documentation for the applicable rules.
Most importantly, CUDs and SUDs should not be treated as additive discounts on the same usage. SUDs do not apply to resource usage already covered by CUDs. Google’s sustained use discount guidance explains how SUD eligibility interacts with other pricing mechanisms.
Choose your CUD type
The right CUD depends on how stable the underlying workload is and how much flexibility you need.Stable Compute Engine workloads
Use a resource-based CUD when you have a durable Compute Engine baseline and the machine configuration and region are unlikely to change.This approach can provide deeper discounts, but the commitment is tied more closely to the resources you specify. It is therefore better suited to workloads with predictable resource requirements.
For the exact resource requirements and eligible configurations, refer to Google Cloud’s committed use discounts documentation.
Variable Compute Engine workloads
Consider a Compute Flexible CUD when your baseline is predictable but the underlying resources may change.Flexible commitments can cover eligible spend across Compute Engine, GKE, and Cloud Run. A single flexible commitment can therefore provide more room to move workloads between supported compute services and resources than a resource-based commitment.
GKE Standard
For GKE Standard, the commitment strategy depends on how the cluster consumes compute.Standard node pools use underlying Compute Engine resources. Applicable Compute Engine resource-based CUDs can therefore cover eligible node compute. Eligible GKE Standard usage can also be covered through Compute Flexible CUDs.
See Google Cloud’s GKE committed use discounts documentation for current GKE eligibility and commitment details.
GKE Autopilot
Autopilot uses a different billing model. Autopilot CUDs are no longer available for new purchases. Google Cloud directs customers toward Compute Flexible CUDs for eligible Autopilot usage.This makes Compute Flexible CUDs the relevant commitment model when predictable Autopilot compute spend is part of the baseline.
The Google Cloud GKE CUD documentation provides the current guidance for Autopilot commitment coverage.
Cloud Run
Cloud Run also fits into the Compute Flexible CUD model for eligible usage. However, discount rates vary by billing model and SKU, so teams should verify the applicable Cloud Run pricing rather than assuming the GKE rate applies.See Google Cloud’s Cloud Run committed use discounts documentation for current eligibility and pricing details.
Teams running a mix of Compute Engine, GKE, and Cloud Run workloads may therefore be able to use a unified flexible commitment rather than managing completely separate compute commitments.
Spot VMs
Spot VMs are a non-commitment option for workloads that can tolerate interruption. They offer variable pricing and discounts of up to 91% off on-demand prices for many machine types, GPUs, TPUs, and Local SSDs, but Google Cloud can preempt them at any time when it needs capacity.Google Cloud’s Spot VMs documentation explains the pricing model, preemption behavior, and supported resources.
Use Spot VMs for fault-tolerant work such as batch processing, CI runners, distributed processing, and workloads that can checkpoint, retry, or be recreated. Do not rely on them as the sole capacity for workloads that cannot withstand interruption.
Spot prices can change, and availability and preemption rates can vary by machine type and zone. Before deploying Spot capacity, check the applicable current price, availability, and historical preemption data for the machine type and location you plan to use.
Cloud SQL and BigQuery
Cloud SQL and BigQuery have their own service-specific commitment programs. They should therefore be evaluated separately from Compute Engine resource-based CUDs and Compute Flexible CUDs.The decision should start with the actual service generating the spend, the stability of that spend, and the commitment model available for that service.
Which savings lever to use first
Start with the least risky optimization and move toward longer commitments only after the baseline is understood.| Your situation | Best first lever | Why |
|---|---|---|
| Oversized or idle VMs | Rightsize first | Prevents committing to waste |
| Eligible steady Compute Engine usage | SUD | Automatic savings without a commitment |
| Stable Compute Engine baseline | Resource-based CUD | Deeper discount for predictable resources |
| Variable but predictable compute spend | Compute Flexible CUD | More flexibility across eligible compute |
| GKE Standard node baseline | Resource-based or Compute Flexible CUD | Depends on resource stability and coverage needs |
| GKE Autopilot baseline | Compute Flexible CUD | Current model for eligible Autopilot commitment spend |
| Heavy BigQuery usage | BigQuery CUD | Service-specific commitment |
| Predictable Cloud SQL fleet | Cloud SQL CUD | Service-specific commitment |
| Fault-tolerant batch jobs | Spot VMs | Lower-cost interruptible capacity |
A commitment sized to an over-provisioned baseline does not make the waste disappear. It locks that baseline into a discounted rate.
Google Cloud’s CUD recommendations can help identify spending and usage patterns and provide recommendations for commitment levels.
Where your GCP spend concentrates
Most cloud bills cluster around a few major services, and each has different optimization mechanics.Compute Engine
Compute Engine is often a major infrastructure cost for GCP workloads.Resource-based CUDs are useful when the baseline is stable enough to commit to specific resources and a region. Compute Flexible CUDs provide a broader model for eligible compute spending when workloads may move across supported services or configurations.
Do not automatically commit the entire Compute Engine bill. Separate:
- Stable baseline usage
- Seasonal or growth-related usage
- Bursty usage
- Idle or oversized resources
- Workloads that may migrate or change machine family
Cloud SQL
Cloud SQL does not use Compute Engine SUDs. It has its own committed use discount model for eligible Cloud SQL usage. When evaluating Cloud SQL commitments, account for the actual production architecture rather than looking only at the primary instance. High-availability configurations can materially change the amount of CPU and memory capacity being consumed.See our Cloud SQL pricing guide for a deeper look at how High Availability changes the cost calculation.
BigQuery
BigQuery has its own spend-based CUD program. It can make sense when analytics consumption is predictable enough to support a commitment.However, do not treat every BigQuery workload as a commitment candidate. Separate recurring baseline usage from highly variable projects, experiments, and temporary analytics spikes before determining the committed amount.
See our BigQuery committed use discounts guide for a deeper look at sizing.
GKE
GKE needs special attention because Standard and Autopilot use different compute models.GKE Standard node compute can use applicable Compute Engine resource-based CUDs, while eligible GKE Standard and Autopilot usage can use Compute Flexible CUDs.
This means teams should identify the GKE operating mode before selecting a commitment strategy.
Autopilot CUDs are no longer available for new purchases; eligible Autopilot usage can instead use Compute Flexible CUDs.
For current eligibility rules, see Google Cloud’s GKE committed use discounts documentation.
Cloud Storage
Cloud Storage does not use Compute CUDs. Storage cost optimization therefore relies on storage class selection, lifecycle management, retention policies, and reducing unnecessary stored data.Do not evaluate a storage-class migration on storage price alone. Retrieval, operations, and network costs can change the overall economics.
An illustrative CUD savings example
Illustrative only. Your rates will vary by machine family, region, commitment type, and coverage.
If the same usage would otherwise qualify for a sustained use discount, the incremental savings from replacing that SUD with a CUD will be lower because the discounts do not stack on the same usage.
For example, if eligible usage would otherwise receive a 30% SUD, the correct comparison is between the CUD-discounted price and the already-SUD-discounted baseline—not between the CUD and the full undiscounted price.
That distinction matters when evaluating whether a commitment is actually worth taking.
The remaining 20% deliberately stays outside the commitment so a burst or temporary growth does not force the team to pay for committed capacity it may not use.
Common commitment mistakes
Committing before rightsizing
A discounted VM is still expensive if the workload does not need the capacity.Rightsize first, then calculate the baseline you are actually willing to commit to.
Assuming CUDs and SUDs stack
They do not provide two independent discounts on the same covered usage.When comparing a CUD against the status quo, account for any SUD that would otherwise apply. See Google Cloud’s sustained use discount guidance for the current rules.
Treating peak usage as the baseline
Peak usage is not the same as committed usage.If demand regularly moves between $30,000 and $50,000 per month, committing the full $50,000 simply because it is the highest observed point increases underutilization risk.
Ignoring GKE mode
GKE Standard and Autopilot do not have identical commitment mechanics.Standard node compute can use applicable Compute Engine resource-based CUDs, while eligible GKE Standard and Autopilot usage can use Compute Flexible CUDs.
Committing during a migration
A migration can make historical usage a poor predictor of future usage.If machine families, regions, architectures, or services are about to change, wait until the new baseline is sufficiently understood before making a long-term commitment.
How we handle commitments at Usage.ai
Native GCP commitments cannot be cancelled after purchase, which makes accurate baseline analysis important before committing. Google Cloud states that after purchasing a CUD, you cannot cancel the commitment and remain responsible for the commitment fee throughout the applicable term.With Flex Insured Commitments, teams can capture commitment-based savings while reducing the risk of unused capacity. Our Flex Commitments reduce cloud spend by 30–50% on average across covered workloads, while qualifying commitments include cashback protection for eligible underutilization.
For GCP teams, we analyze billing data to identify stable usage that may be suitable for commitments, help determine an appropriate commitment level, and support eligible commitment actions.
This helps teams capture more of the savings available through GCP commitments without taking on the traditional risk of unused capacity.
Automate Compute Engine CUDs and flexible commitments with Usage.ai and get cashback protection without the risk of long-term commitments.
Frequently asked questions
Does DynamoDB Streams consume my table's read capacity?
No. Streams reads are billed in their own Streams read request units, which are entirely separate from the read capacity units your application uses against the table. Enabling and reading a stream does not consume or reduce your table's provisioned or on-demand read capacity.
Do Lambda functions always read Streams for free?
Almost always. Standard serverless Lambda functions triggered off a stream read for free, and Global Tables replication reads are free too. The one exception is a function running on Lambda Managed Instances, which is billed at the standard $0.02 per 100,000 rate. You still pay separately for Lambda requests and duration in every case.
Do TTL deletions generate stream records?
Yes. When TTL expires an item, DynamoDB writes a delete record to the stream if Streams is enabled. It carries a userIdentity field marking it as a service-initiated delete. Streams does not charge per record, but for a custom consumer a high TTL expiry rate can still raise cost by increasing GetRecords calls. You can filter to invoke your function only on those records, or drop all deletes by filtering event type.
How long are Streams records retained, and can I extend it?
Records are retained for 24 hours and then deleted automatically. There is no setting to extend that window. If a consumer falls more than 24 hours behind, it loses records. When you need a longer replay window, route changes to Kinesis Data Streams, which supports retention up to 365 days.
How do I monitor my Streams read volume for cost?
Use the CloudWatch SuccessfulRequestLatency metric with the SampleCount statistic, which reports the number of GetRecords calls per stream. Because it counts free reads (Lambda, Global Tables) alongside billable ones, subtract those to isolate charged reads, then alarm on the result. DynamoDB Streams does not support cost allocation tags, so CloudWatch is the primary native visibility path.