New See exactly what you're overpaying AWS in under 60 seconds. Try the Calculator for free

GCP May 2026 Updates: Gemini, AI Commitments, and Key Cloud Changes

Google Cloud’s May 2026 updates expand Gemini, AI capabilities, and cloud commitments, with key changes affecting performance, flexibility, and costs.
Updated August 12, 2026
15 min read
GCP in May 2026: Gemini 3.5 Flash Lands at Google I/O, 2.0 Models Die June 1, and Airflow 3.1 Reaches GA
In this article
Key takeaways
1
Gemini 2.0 Flash and Flash-Lite were scheduled to shut down June 1. Teams still using affected model IDs needed to plan migration and testing.
2
Gemini 3.5 Flash launched May 19. Its Gemini Developer API Standard price was $1.50 per million input tokens and $9 per million output tokens.
3
Qualifying Agent Platform online inference can use Compute Engine reservations. Applicable discounts affect reserved Compute Engine resources.
For FinOps teams, May’s most important GCP changes were concentrated around AI model lifecycle, inference infrastructure, and migration economics.

This roundup covers Google Cloud announcements and product changes from May 2026, with the June 1 Gemini shutdown included because it was a directly relevant deadline for May planning.

Gemini 3.5 Flash: benchmark before changing models

Gemini 3.5 Flash launched on May 19, 2026, at Google I/O. Google positioned it for coding, reasoning, and agentic workloads.

For the Gemini Developer API, Standard paid pricing was $1.50 per million input tokens and $9 per million output tokens. Other pricing modes can differ, so this rate should not be treated as a universal Vertex AI or Agent Platform price. See Google Gemini API pricing.

Don’t compare models by token price alone. Compare what it costs to get the same job done. A cheaper model may use more tokens, take longer, produce weaker results, or require more retries.

Before switching models, run a sample of real workloads through both the current model and the candidate model. Compare total cost, response quality, latency, and retries. The most useful metric is cost per successful task, not just the price per million tokens.

For Gemini workloads, also check whether caching or thinking tokens affect the final cost.

Gemini 2.0 Flash and Flash-Lite: June 1 shutdown

Google announced on February 18 that Gemini 2.0 Flash, Gemini 2.0 Flash-001, Gemini 2.0 Flash-Lite, and Gemini 2.0 Flash-Lite-001 would shut down on June 1, 2026. This made model inventory a May planning priority.

Google’s current deprecation guidance recommends Gemini 3.6 Flash for Gemini 2.0 Flash and Gemini 3.1 Flash-Lite for Gemini 2.0 Flash-Lite.

Because migration guidance can change, teams should verify the target model and pricing in the product surface they use before making the change.

Start by finding every application that uses the retiring model. Check application code, configuration files, deployment settings, API clients, and monitoring tools for the old model ID.

Then compare how much those applications currently spend and how many input and output tokens they use. Test the replacement model with the same workloads before switching.

Do not assume the replacement will cost the same. A new model can have different prices and may use a different number of tokens to complete the same task.

Related read: AWS in May 2026 Updates

Agent Platform: review reservation-backed coverage

Google documents Compute Engine reservations for qualifying Agent Platform online inference. Reserved resources are charged at the applicable Compute Engine rate, including applicable discounts. See Google Agent Platform reservation documentation.

Agent Platform costs are not all covered by Compute Engine commitments. When eligible inference workloads use reservation-backed Compute Engine resources, the Compute Engine portion can receive the applicable commitment benefit. Separate Agent Platform fees still need to be paid and evaluated separately.

Check eligibility before counting this usage toward your commitment baseline. Confirm that the workload, VM type, region, zone, reservation setup, and sharing configuration meet Google’s requirements.

If you already have unused Compute Engine commitment capacity, first check whether eligible Agent Platform infrastructure can use it. Do not buy additional commitments just because you run Agent Platform. Establish that the underlying usage is stable, eligible, and likely to continue before increasing your commitment.

App Engine Migration Hub: useful, but Preview

On May 21, Google announced that the App Engine Migration Hub can migrate eligible App Engine standard-environment services to Cloud Run and provide cost-saving recommendations. The migration workflow is Preview.

The documented workflow applies to second-generation runtimes that do not use App Engine legacy bundled services. It is therefore not a universal migration path for every App Engine application.

For FinOps teams, the value is the ability to evaluate a possible Cloud Run migration before committing to the move. Treat estimated savings as an input to the business case, not guaranteed savings.

Cloud Run can scale to zero, but actual economics depend on configuration and workload behavior. Check minimum instances, CPU and memory requirements, billing mode, networking, startup behavior, and dependent services before comparing the two environments.

Knowledge Catalog data products reach GA

Data products in Knowledge Catalog reached general availability on May 25. The release included approval workflows for data-product consumption and other capabilities for publishing, accessing, and managing curated data products.

This is primarily a governance and data-management update rather than a direct pricing change. For FinOps teams, the potential value is indirect: stronger governance can improve visibility into ownership, access, and consumption, but GA status itself does not reduce cloud spend.

GKE May change: backend authenticated TLS

On May 29, GKE Gateway added backend authenticated TLS for Gateway-originated connections to supported Pods or InferencePools and GatewayClasses.

This is a security and networking capability, not a direct cost optimization. It should not be presented as mutual TLS: backend authenticated TLS lets the Gateway verify the backend’s identity, while mTLS involves bidirectional identity verification.

Airflow 3: April GA, May build update

Airflow 3 should not be described as a May GA event. Google’s Managed Service for Apache Airflow release notes show a May 14, 2026 release, so the May takeaway is release and upgrade planning rather than a new GA milestone.

Teams running Managed Service for Apache Airflow should review the May build, compatibility considerations, and available upgrade paths before changing production environments.

Related read: Azure in May 2026 Updates

A Checklist for FinOps teams do

Audit Gemini model IDs.

Identify affected Gemini 2.0 references and confirm supported replacement models.

Benchmark before changing models.

Compare cost per successful task, not just token price.

Review Agent Platform coverage.

Separate eligible Compute Engine infrastructure from other Agent Platform charges and verify reservation requirements.

Validate App Engine migration economics.

Check runtime eligibility and configuration before relying on Cloud Run savings estimates.

Separate operational releases from savings opportunities.

Airflow, GKE, and Knowledge Catalog changes can improve operations or governance without directly lowering spend.

How Usage.ai fits

We help teams identify and manage eligible GCP commitment opportunities based on actual usage patterns.

With Flex Insured Commitments, teams can capture commitment-based savings while reducing the risk of unused capacity. Our Flex Commitments reduce cloud spend by 30–50% on average across covered workloads, while qualifying commitments include cashback protection for eligible underutilization.

For GCP teams, we analyze billing data to identify stable usage that may be suitable for commitments, help determine an appropriate commitment level, and support eligible commitment actions.

This helps teams capture more of the savings available through GCP commitments without taking on the traditional risk of unused capacity.
Evaluate with your own data
Run a Free Savings Analysis.

Connect in 15 minutes. No contracts, no infrastructure changes. See your savings before committing.

Frequently asked questions

What are the most important GCP May 2026 updates for FinOps teams?

The most relevant changes are the Gemini 2.0 shutdown, Gemini 3.5 Flash launch, reservation-backed Agent Platform inference, and the App Engine Migration Hub Preview. The main actions are model migration, workload benchmarking, commitment-coverage review, and migration-cost validation.

When did Gemini 2.0 Flash and Flash-Lite shut down?

They were scheduled to shut down on June 1, 2026. Teams using affected model IDs needed to migrate and test before that deadline.

What is the current replacement for Gemini 2.0 Flash?

Google’s current Gemini API deprecation guidance recommends Gemini 3.6 Flash for Gemini 2.0 Flash and Gemini 3.1 Flash-Lite for Gemini 2.0 Flash-Lite. Verify the target model in the product surface you use before migrating.

Can CUDs reduce Agent Platform inference costs?

Applicable Compute Engine discounts can reduce the cost of qualifying reserved Compute Engine resources used for inference. Teams must separately evaluate other Agent Platform charges and confirm reservation eligibility.

Can App Engine Migration Hub guarantee Cloud Run savings?

No. Migration Hub provides cost estimates and recommendations for eligible workloads, but the workflow is Preview. Actual savings depend on application configuration, traffic, billing mode, resources, networking, and other workload characteristics.

Share
Facebook
X
LinkedIn
Reddit
Cut cloud cost with automation
Latest from our blogs