Cloud & Infrastructure News: Aurora PostgreSQL Queries Iceberg and Parquet Directly, S3 Vectors Gets Metadata Pre-Filtering, S3 Tables Add Iceberg V3 Types, 2026-10-03
cloud

Cloud & Infrastructure News: Aurora PostgreSQL Queries Iceberg and Parquet Directly, S3 Vectors Gets Metadata Pre-Filtering, S3 Tables Add Iceberg V3 Types, 2026-10-03

4 min read

Amazon Aurora PostgreSQL Can Now Query Iceberg and Parquet in S3 Directly

Aurora PostgreSQL has a new aurora_analytics extension that embeds a DuckDB engine. It lets you query Apache Iceberg tables and Parquet files in S3 from the same database that holds your transactional data, with no ETL job and no separate warehouse. You run CREATE EXTENSION aurora_analytics;, then create a foreign table on aurora_analytics_server that points at an S3 location and format. The schema is inferred from the Parquet/Iceberg metadata, so you can write CREATE FOREIGN TABLE transaction_history () ... OPTIONS (location 's3://.../file.parquet', format 'parquet') with an empty column list.

The typical use case is a single SQL query that joins recent rows in Aurora with years of history archived to the lake. The engine pushes down predicates and prunes columns, caches frequently read data locally on the instance, and exposes aurora_analytics_stat_statements() so you can see rows scanned and S3 bytes read per query.

It requires Aurora PostgreSQL 17.11+ or 18.6+ and is available in all commercial regions and GovCloud (US). There is no feature charge: you pay for the extra Aurora compute you use and standard S3 GET requests. For teams already offloading cold data to Iceberg, this can replace a federated-query setup or a nightly copy back into Postgres for reporting.

Read more — AWS News Blog


Amazon S3 Vectors Adds Metadata Pre-Filtering for Higher Recall on Filtered Queries

S3 Vectors indexes now support an ENHANCED mode that evaluates the metadata filter first and runs the similarity search only over vectors that match. The existing behavior, now called CLASSIC, evaluates the filter during the vector search, which can return far fewer than topK results when the filter is very selective. AWS says pre-filtering returns up to 5x more matching vectors on highly selective filters. In its single-tenant example, a query went from 2 results to the full 10.

Each vector can carry up to 2 KB of filterable metadata, and a query can include up to 100 filter constraints using a compact JSON syntax with operators such as $and, $or, $gt and $startsWith. New indexes choose a mode at CreateIndex, and existing indexes can be switched with the new UpdateIndexMode operation.

This matters most for multi-tenant RAG, where every query is scoped by tenant, document status or date range. Under post-filtering, small tenants in large indexes routinely got empty or truncated results. The feature costs nothing extra and is available in every commercial region where S3 Vectors runs, plus the China regions.

Read more — AWS News Blog


Amazon S3 Tables Now Support All Apache Iceberg V3 Data Types

S3 Tables now covers the full set of data types added in Apache Iceberg V3:

  • variant for semi-structured data, stored in columnar form and automatically shredded into hidden columns so queries like variant_get() don't parse JSON at runtime
  • Nanosecond timestamp / timestamptz for high-precision event data
  • geometry and geography for native geospatial columns
  • unknown for columns whose type isn't known yet

V3 tables also get deletion vectors, which replace positional delete files and make row-level deletes and updates cheaper, and row lineage through automatic _row_id and _last_updated_sequence_number columns, which helps with incremental processing and CDC. Existing V2 tables can be upgraded in place with ALTER TABLE without rewriting data.

Using the new types requires a Spark 4.0+ based engine, such as AWS Glue 6.0+ or Amazon EMR 8.1+, and they work across Redshift, EMR and Glue. There is no additional charge. Together with the Aurora integration above, this means you can write variant-typed event data from Spark and read it from Postgres in the same lake.

Read more — AWS News Blog


Stanislav Lentsov

Written by

Stanislav Lentsov

Software Architect

You May Also Enjoy