AWS just changed how Aurora can access data lake data
On September 30, 2026, AWS announced that Amazon Aurora PostgreSQL can directly query Apache Iceberg and Parquet data stored in data lakes.Rather than copying data into Aurora first, applications can create PostgreSQL foreign tables that reference data in Amazon S3, S3 Tables, or AWS Glue Data Catalog and query it alongside native Aurora tables.
AWS embedded DuckDB in Aurora PostgreSQL to execute the analytical portion of these queries.
Read the full technical announcement in the AWS News Blog or AWS’s shorter What’s New announcement.
The capability is generally available in Aurora PostgreSQL 17.11 and 18.6 and includes Aurora Serverless v2.
This guide helps you assess whether direct data lake querying lowers total cost, shifts costs from ETL and Aurora storage to compute and S3 requests, or increases overall spend.
What changed at a glance
| Announced | September 30, 2026 |
| Service | Amazon Aurora PostgreSQL |
| Supported versions | 17.11+, 18.6+ |
| Formats | Apache Iceberg and Apache Parquet |
| Query engine | DuckDB embedded in Aurora PostgreSQL |
| Data sources | Amazon S3, S3 Tables, AWS Glue Data Catalog, and supported federated catalogs |
| Additional feature charge | None |
| Costs still incurred | Incremental Aurora compute and Amazon S3 requests |
The feature is technically useful, but CTOs and FinOps teams face a more important question:
The bigger change is where your data needs to live
Consider an application that keeps current customer transactions in Aurora while several years of historical records remain in Iceberg or Parquet in a data lake.If the application needs both datasets, the team may maintain a pipeline that copies the needed historical data into Aurora.
That can create another copy of the data, infrastructure to move it, and engineering work to keep everything synchronized.
Direct querying changes that architecture.
Before
Now, for suitable workloads
That does not mean ETL is suddenly unnecessary.
Data pipelines will still make sense for many workloads, transformations, and latency requirements. AWS also supports materializing Iceberg or Parquet data into native Aurora tables when applications need a different access pattern.
The important change is that teams now have another choice about which data needs to live inside Aurora at all.
For teams evaluating that decision primarily through cost, our RDS vs. Aurora cost comparison provides broader context on the components that shape Aurora total cost.
“No additional charge” does not mean “no additional cost”
This is where the announcement becomes a FinOps question.AWS says the capability itself has no additional charge. Customers still pay for the incremental Aurora compute used by queries and Amazon S3 request costs for reading data lake files.
The feature can reshape the AWS bill rather than simply remove a cost.
| You may reduce | You may increase |
|---|---|
| Data duplicated inside Aurora | Aurora query compute |
| Infrastructure used to copy data into Aurora | Amazon S3 request activity |
| Operational work maintaining some data pipelines | Compute required for analytical workloads |
| Aurora data footprint for data that remains in S3 | Serverless capacity as query demand changes |
Occasional queries of historical customer records differ greatly from repeated scans of large datasets with concurrent analytical queries.
How to evaluate the cost tradeoff
Direct querying is more likely to make sense when applications need selective or occasional access to historical data and can tolerate reading it from the lake.So, before moving data out of Aurora, teams need to test actual query patterns. Assess how much data each query scans, filter selectivity, how often the same data is requested, query concurrency, and latency requirements.
Aurora uses techniques such as predicate pushdown, column pruning, and caching to reduce unnecessary reads, but workload shape still matters.
For latency-sensitive query patterns, AWS supports materializing Iceberg or Parquet data into native Aurora tables.
Test the cost decision rather than assuming it. Then, compare a representative period before and after the change.
For many steady workloads, two to four weeks may capture normal usage patterns, while seasonal or highly variable workloads may require a longer window.
Measure what disappears from the old architecture:
duplicated Aurora storage
reverse ETL or data-movement infrastructure
operational effort associated with those pipelines
Aurora compute or Serverless v2 ACU consumption
S3 requests and bytes read from S3
reader capacity, if additional readers are used
query latency and concurrency
frequency and volume of analytical queries
Pay particular attention to Aurora compute
This is especially important if you use Aurora Serverless v2.Serverless v2 automatically adjusts database capacity using Aurora Capacity Units, or ACUs. If direct data lake queries increase database demand, capacity consumption can change as well.
That does not necessarily mean the feature will make Serverless v2 more expensive. It means you need to measure the resulting consumption.
For a deeper look at ACUs, scaling, and pricing, see our Aurora Serverless v2 pricing and cost guide.
For provisioned Aurora clusters, teams should similarly determine whether new queries affect CPU, memory, concurrency, or the need for additional reader capacity.
Establish the new baseline before increasing commitments
There is another easy-to-miss consequence.AWS Database Savings Plans offer discounted rates in exchange for a one-year commitment to a consistent dollar-per-hour amount of eligible database usage, including Aurora.
They differ from Aurora Reserved DB Instances, which can run for one or three years and are tied to Aurora instance configuration. The two discounts cannot apply to the same usage.
If your Aurora architecture changes, the usage pattern you want to commit against may change too.
Consider three possible outcomes:
Scenario 1: You remove infrastructure that copied historical data into Aurora while Aurora compute remains relatively stable.
Scenario 2: You remove the pipeline but need more Aurora compute or reader capacity to support analytical queries.
Scenario 3: Your Aurora Serverless v2 workload begins scaling differently because analytical queries are now part of its demand.
Teams should assess:
Aurora compute consumption
Serverless v2 ACU consumption, where applicable
S3 request activity
The amount of data stored in Aurora
Infrastructure retired from the previous architecture
The frequency and intensity of new queries
Our AWS Database Savings Plans guide covers eligibility, coverage, and commitment sizing in more detail.
If an architecture change materially changes your durable Aurora consumption, your commitment baseline should change with it.
What this means for CTOs and FinOps teams
Aurora PostgreSQL’s new Iceberg and Parquet support gives teams another choice about where historical data lives and where the work to query it happens.For engineering teams, this can simplify applications that need both operational and historical data, while for FinOps teams, the task is to quantify the tradeoff.
Measure what disappears from the old architecture.
Measure the Aurora and S3 consumption that takes its place.
Then establish the new database baseline before deciding how much of that spend should be committed.
Changing your database architecture?
Run our Savings Test to assess which AWS spend is ready for additional commitment coverage. For eligible Flex Insured Commitments, we can help you access up to 57% savings associated with a three-year AWS commitment while reducing long-term commitment exposure.
We also offer cashback protection when a covered commitment costs more than equivalent On-Demand usage.
Review Aurora PostgreSQL, Iceberg, and Parquet query patterns before changing your data architecture.
Frequently asked questions
Can Aurora PostgreSQL query Iceberg and Parquet data without ETL?
Yes. Aurora PostgreSQL can query Iceberg and Parquet data directly through foreign tables, so teams do not always need to copy that data into Aurora first.
Does querying Iceberg and Parquet from Aurora cost extra?
There is no separate feature fee. You still pay for the Aurora compute used by the queries and the Amazon S3 requests generated.
When should you query data directly instead of materializing it in Aurora?
Direct querying can fit selective or occasional access to historical data. Materialization may be better when you need consistently low latency or repeatedly access the same data.