This roundup covers Google Cloud announcements and product changes from May 2026, with the June 1 Gemini shutdown included because it was a directly relevant deadline for May planning.
Gemini 3.5 Flash: benchmark before changing models
For the Gemini Developer API, Standard paid pricing was $1.50 per million input tokens and $9 per million output tokens. Other pricing modes can differ, so this rate should not be treated as a universal Vertex AI or Agent Platform price. See Google Gemini API pricing.
Don’t compare models by token price alone. Compare what it costs to get the same job done. A cheaper model may use more tokens, take longer, produce weaker results, or require more retries.
Before switching models, run a sample of real workloads through both the current model and the candidate model. Compare total cost, response quality, latency, and retries. The most useful metric is cost per successful task, not just the price per million tokens.
For Gemini workloads, also check whether caching or thinking tokens affect the final cost.
Gemini 2.0 Flash and Flash-Lite: June 1 shutdown
Google’s current deprecation guidance recommends Gemini 3.6 Flash for Gemini 2.0 Flash and Gemini 3.1 Flash-Lite for Gemini 2.0 Flash-Lite.
Because migration guidance can change, teams should verify the target model and pricing in the product surface they use before making the change.
Start by finding every application that uses the retiring model. Check application code, configuration files, deployment settings, API clients, and monitoring tools for the old model ID.
Then compare how much those applications currently spend and how many input and output tokens they use. Test the replacement model with the same workloads before switching.
Do not assume the replacement will cost the same. A new model can have different prices and may use a different number of tokens to complete the same task.
Related read: AWS in May 2026 Updates
Agent Platform: review reservation-backed coverage
Agent Platform costs are not all covered by Compute Engine commitments. When eligible inference workloads use reservation-backed Compute Engine resources, the Compute Engine portion can receive the applicable commitment benefit. Separate Agent Platform fees still need to be paid and evaluated separately.
Check eligibility before counting this usage toward your commitment baseline. Confirm that the workload, VM type, region, zone, reservation setup, and sharing configuration meet Google’s requirements.
If you already have unused Compute Engine commitment capacity, first check whether eligible Agent Platform infrastructure can use it. Do not buy additional commitments just because you run Agent Platform. Establish that the underlying usage is stable, eligible, and likely to continue before increasing your commitment.
App Engine Migration Hub: useful, but Preview
The documented workflow applies to second-generation runtimes that do not use App Engine legacy bundled services. It is therefore not a universal migration path for every App Engine application.
For FinOps teams, the value is the ability to evaluate a possible Cloud Run migration before committing to the move. Treat estimated savings as an input to the business case, not guaranteed savings.
Cloud Run can scale to zero, but actual economics depend on configuration and workload behavior. Check minimum instances, CPU and memory requirements, billing mode, networking, startup behavior, and dependent services before comparing the two environments.
Knowledge Catalog data products reach GA
This is primarily a governance and data-management update rather than a direct pricing change. For FinOps teams, the potential value is indirect: stronger governance can improve visibility into ownership, access, and consumption, but GA status itself does not reduce cloud spend.
GKE May change: backend authenticated TLS
This is a security and networking capability, not a direct cost optimization. It should not be presented as mutual TLS: backend authenticated TLS lets the Gateway verify the backend’s identity, while mTLS involves bidirectional identity verification.
Airflow 3: April GA, May build update
Teams running Managed Service for Apache Airflow should review the May build, compatibility considerations, and available upgrade paths before changing production environments.
Related read: Azure in May 2026 Updates
A Checklist for FinOps teams do
Identify affected Gemini 2.0 references and confirm supported replacement models.
Compare cost per successful task, not just token price.
Separate eligible Compute Engine infrastructure from other Agent Platform charges and verify reservation requirements.
Check runtime eligibility and configuration before relying on Cloud Run savings estimates.
Airflow, GKE, and Knowledge Catalog changes can improve operations or governance without directly lowering spend.
How Usage.ai fits
With Flex Insured Commitments, teams can capture commitment-based savings while reducing the risk of unused capacity. Our Flex Commitments reduce cloud spend by 30–50% on average across covered workloads, while qualifying commitments include cashback protection for eligible underutilization.
For GCP teams, we analyze billing data to identify stable usage that may be suitable for commitments, help determine an appropriate commitment level, and support eligible commitment actions.
This helps teams capture more of the savings available through GCP commitments without taking on the traditional risk of unused capacity.
Connect in 15 minutes. No contracts, no infrastructure changes. See your savings before committing.
Frequently asked questions
What are the most important GCP May 2026 updates for FinOps teams?
The most relevant changes are the Gemini 2.0 shutdown, Gemini 3.5 Flash launch, reservation-backed Agent Platform inference, and the App Engine Migration Hub Preview. The main actions are model migration, workload benchmarking, commitment-coverage review, and migration-cost validation.
When did Gemini 2.0 Flash and Flash-Lite shut down?
They were scheduled to shut down on June 1, 2026. Teams using affected model IDs needed to migrate and test before that deadline.
What is the current replacement for Gemini 2.0 Flash?
Google’s current Gemini API deprecation guidance recommends Gemini 3.6 Flash for Gemini 2.0 Flash and Gemini 3.1 Flash-Lite for Gemini 2.0 Flash-Lite. Verify the target model in the product surface you use before migrating.
Can CUDs reduce Agent Platform inference costs?
Applicable Compute Engine discounts can reduce the cost of qualifying reserved Compute Engine resources used for inference. Teams must separately evaluate other Agent Platform charges and confirm reservation eligibility.
Can App Engine Migration Hub guarantee Cloud Run savings?
No. Migration Hub provides cost estimates and recommendations for eligible workloads, but the workflow is Preview. Actual savings depend on application configuration, traffic, billing mode, resources, networking, and other workload characteristics.