New See exactly what you're overpaying AWS in under 60 seconds. Try the Calculator for free →

AWS Announces Aurora PostgreSQL Iceberg and Parquet Support: What Direct Queries Change for Cost

What AWS’s new data lake query support means for Aurora costs and commitments.
Updated October 5, 2026
19 min read
AWS Announces Aurora PostgreSQL Iceberg and Parquet Support: What Direct Queries Change for Cost
In this article
Key takeaways
1
Aurora PostgreSQL can now query Apache Iceberg and Parquet data directly alongside operational database data.
2
Account for autoscale, active regions, reservation scope, and existing coverage before purchasing more.
3
If this change affects your steady-state Aurora consumption, establish a new baseline before making additional database commitments.

AWS just changed how Aurora can access data lake data

On September 30, 2026, AWS announced that Amazon Aurora PostgreSQL can directly query Apache Iceberg and Parquet data stored in data lakes.

Rather than copying data into Aurora first, applications can create PostgreSQL foreign tables that reference data in Amazon S3, S3 Tables, or AWS Glue Data Catalog and query it alongside native Aurora tables.

AWS embedded DuckDB in Aurora PostgreSQL to execute the analytical portion of these queries.

Read the full technical announcement in the AWS News Blog or AWS’s shorter What’s New announcement.

The capability is generally available in Aurora PostgreSQL 17.11 and 18.6 and includes Aurora Serverless v2.

This guide helps you assess whether direct data lake querying lowers total cost, shifts costs from ETL and Aurora storage to compute and S3 requests, or increases overall spend.

What changed at a glance

Announced September 30, 2026
Service Amazon Aurora PostgreSQL
Supported versions 17.11+, 18.6+
Formats Apache Iceberg and Apache Parquet
Query engine DuckDB embedded in Aurora PostgreSQL
Data sources Amazon S3, S3 Tables, AWS Glue Data Catalog, and supported federated catalogs
Additional feature charge None
Costs still incurred Incremental Aurora compute and Amazon S3 requests
Source: AWS documentation on querying Apache Iceberg and Parquet data directly in Aurora PostgreSQL

The feature is technically useful, but CTOs and FinOps teams face a more important question:
If some data no longer needs to be copied into Aurora, what happens to the cost of the architecture?

The bigger change is where your data needs to live

Consider an application that keeps current customer transactions in Aurora while several years of historical records remain in Iceberg or Parquet in a data lake.

If the application needs both datasets, the team may maintain a pipeline that copies the needed historical data into Aurora.

That can create another copy of the data, infrastructure to move it, and engineering work to keep everything synchronized.

Direct querying changes that architecture.

Before

Data lake → data pipeline → Aurora copy → application

Now, for suitable workloads

Data lake → Aurora PostgreSQL → application
AWS explicitly positions the capability as a way to access data lake data without first moving or duplicating it.

That does not mean ETL is suddenly unnecessary.

Data pipelines will still make sense for many workloads, transformations, and latency requirements. AWS also supports materializing Iceberg or Parquet data into native Aurora tables when applications need a different access pattern.

The important change is that teams now have another choice about which data needs to live inside Aurora at all.

For teams evaluating that decision primarily through cost, our RDS vs. Aurora cost comparison provides broader context on the components that shape Aurora total cost.

“No additional charge” does not mean “no additional cost”

This is where the announcement becomes a FinOps question.

AWS says the capability itself has no additional charge. Customers still pay for the incremental Aurora compute used by queries and Amazon S3 request costs for reading data lake files.

The feature can reshape the AWS bill rather than simply remove a cost.
You may reduce You may increase
Data duplicated inside Aurora Aurora query compute
Infrastructure used to copy data into Aurora Amazon S3 request activity
Operational work maintaining some data pipelines Compute required for analytical workloads
Aurora data footprint for data that remains in S3 Serverless capacity as query demand changes
The outcome depends on the workload.

Occasional queries of historical customer records differ greatly from repeated scans of large datasets with concurrent analytical queries.
You need to ask: Does the infrastructure and data duplication we can remove cost more than the Aurora compute and S3 consumption we add?
Teams should make that calculation before treating the new architecture as a cost optimization.

How to evaluate the cost tradeoff

Direct querying is more likely to make sense when applications need selective or occasional access to historical data and can tolerate reading it from the lake.

So, before moving data out of Aurora, teams need to test actual query patterns. Assess how much data each query scans, filter selectivity, how often the same data is requested, query concurrency, and latency requirements.

Aurora uses techniques such as predicate pushdown, column pruning, and caching to reduce unnecessary reads, but workload shape still matters.

For latency-sensitive query patterns, AWS supports materializing Iceberg or Parquet data into native Aurora tables.

Test the cost decision rather than assuming it. Then, compare a representative period before and after the change. 

For many steady workloads, two to four weeks may capture normal usage patterns, while seasonal or highly variable workloads may require a longer window.

Measure what disappears from the old architecture:

duplicated Aurora storage

reverse ETL or data-movement infrastructure

operational effort associated with those pipelines

Then measure what changes or is added:

Aurora compute or Serverless v2 ACU consumption

S3 requests and bytes read from S3

reader capacity, if additional readers are used

query latency and concurrency

frequency and volume of analytical queries

Use application and database monitoring alongside AWS billing data to validate the before-and-after comparison.

Pay particular attention to Aurora compute

This is especially important if you use Aurora Serverless v2.

Serverless v2 automatically adjusts database capacity using Aurora Capacity Units, or ACUs. If direct data lake queries increase database demand, capacity consumption can change as well.

That does not necessarily mean the feature will make Serverless v2 more expensive. It means you need to measure the resulting consumption.

For a deeper look at ACUs, scaling, and pricing, see our Aurora Serverless v2 pricing and cost guide.

For provisioned Aurora clusters, teams should similarly determine whether new queries affect CPU, memory, concurrency, or the need for additional reader capacity.
Note: Removing a data pipeline does not automatically mean the workload behind that pipeline disappears. Some of that work may now happen somewhere else.

Establish the new baseline before increasing commitments

There is another easy-to-miss consequence.

AWS Database Savings Plans offer discounted rates in exchange for a one-year commitment to a consistent dollar-per-hour amount of eligible database usage, including Aurora.

They differ from Aurora Reserved DB Instances, which can run for one or three years and are tied to Aurora instance configuration. The two discounts cannot apply to the same usage.

If your Aurora architecture changes, the usage pattern you want to commit against may change too.

Consider three possible outcomes:

Scenario 1: You remove infrastructure that copied historical data into Aurora while Aurora compute remains relatively stable.

Scenario 2: You remove the pipeline but need more Aurora compute or reader capacity to support analytical queries.

Scenario 3: Your Aurora Serverless v2 workload begins scaling differently because analytical queries are now part of its demand.

These are not predictions. They illustrate why you need to measure the new environment before making a commitment decision.

Teams should assess:

Aurora compute consumption

Serverless v2 ACU consumption, where applicable

S3 request activity

The amount of data stored in Aurora

Infrastructure retired from the previous architecture

The frequency and intensity of new queries

Then determine which portion of the resulting database spend is durable enough to commit.

Our AWS Database Savings Plans guide covers eligibility, coverage, and commitment sizing in more detail.
The key principle is:
If an architecture change materially changes your durable Aurora consumption, your commitment baseline should change with it.
Otherwise, you risk sizing future commitments against an architecture that no longer reflects how the database is used.

What this means for CTOs and FinOps teams

Aurora PostgreSQL’s new Iceberg and Parquet support gives teams another choice about where historical data lives and where the work to query it happens.

For engineering teams, this can simplify applications that need both operational and historical data, while for FinOps teams, the task is to quantify the tradeoff.
1

Measure what disappears from the old architecture.

2

Measure the Aurora and S3 consumption that takes its place.

3

Then establish the new database baseline before deciding how much of that spend should be committed.

That is how an architectural improvement has a better chance of becoming a cost improvement rather than simply moving spend from one AWS service or resource to another.

Changing your database architecture? 

Run our Savings Test to assess which AWS spend is ready for additional commitment coverage. For eligible Flex Insured Commitments, we can help you access up to 57% savings associated with a three-year AWS commitment while reducing long-term commitment exposure.

We also offer cashback protection when a covered commitment costs more than equivalent On-Demand usage.
AURORA COST REVIEW
Evaluate Direct Query Economics Before You Scale

Review Aurora PostgreSQL, Iceberg, and Parquet query patterns before changing your data architecture.

Frequently asked questions

Can Aurora PostgreSQL query Iceberg and Parquet data without ETL?

Yes. Aurora PostgreSQL can query Iceberg and Parquet data directly through foreign tables, so teams do not always need to copy that data into Aurora first.

Does querying Iceberg and Parquet from Aurora cost extra?

There is no separate feature fee. You still pay for the Aurora compute used by the queries and the Amazon S3 requests generated.

When should you query data directly instead of materializing it in Aurora?

Direct querying can fit selective or occasional access to historical data. Materialization may be better when you need consistently low latency or repeatedly access the same data.

Share
Facebook
X
LinkedIn
Reddit
Cut cloud cost with automation
Latest from our blogs