Amazon Aurora PostgreSQL Can Now Query Iceberg and Parquet in S3 Directly
Aurora PostgreSQL has a new aurora_analytics extension that embeds a DuckDB engine. It lets you query Apache Iceberg tables and Parquet files in S3 from the same database that holds your transactional data, with no ETL job and no separate warehouse. You run CREATE EXTENSION aurora_analytics;, then create a foreign table on aurora_analytics_server that points at an S3 location and format. The schema is inferred from the Parquet/Iceberg metadata, so you can write CREATE FOREIGN TABLE transaction_history () ... OPTIONS (location 's3://.../file.parquet', format 'parquet') with an empty column list.
The typical use case is a single SQL query that joins recent rows in Aurora with years of history archived to the lake. The engine pushes down predicates and prunes columns, caches frequently read data locally on the instance, and exposes aurora_analytics_stat_statements() so you can see rows scanned and S3 bytes read per query.
It requires Aurora PostgreSQL 17.11+ or 18.6+ and is available in all commercial regions and GovCloud (US). There is no feature charge: you pay for the extra Aurora compute you use and standard S3 GET requests. For teams already offloading cold data to Iceberg, this can replace a federated-query setup or a nightly copy back into Postgres for reporting.
Read more — AWS News Blog
Amazon S3 Vectors Adds Metadata Pre-Filtering for Higher Recall on Filtered Queries
S3 Vectors indexes now support an ENHANCED mode that evaluates the metadata filter first and runs the similarity search only over vectors that match. The existing behavior, now called CLASSIC, evaluates the filter during the vector search, which can return far fewer than topK results when the filter is very selective. AWS says pre-filtering returns up to 5x more matching vectors on highly selective filters. In its single-tenant example, a query went from 2 results to the full 10.
Each vector can carry up to 2 KB of filterable metadata, and a query can include up to 100 filter constraints using a compact JSON syntax with operators such as $and, $or, $gt and $startsWith. New indexes choose a mode at CreateIndex, and existing indexes can be switched with the new UpdateIndexMode operation.
This matters most for multi-tenant RAG, where every query is scoped by tenant, document status or date range. Under post-filtering, small tenants in large indexes routinely got empty or truncated results. The feature costs nothing extra and is available in every commercial region where S3 Vectors runs, plus the China regions.
Read more — AWS News Blog
Amazon S3 Tables Now Support All Apache Iceberg V3 Data Types
S3 Tables now covers the full set of data types added in Apache Iceberg V3:
variantfor semi-structured data, stored in columnar form and automatically shredded into hidden columns so queries likevariant_get()don't parse JSON at runtime- Nanosecond
timestamp/timestamptzfor high-precision event data geometryandgeographyfor native geospatial columnsunknownfor columns whose type isn't known yet
V3 tables also get deletion vectors, which replace positional delete files and make row-level deletes and updates cheaper, and row lineage through automatic _row_id and _last_updated_sequence_number columns, which helps with incremental processing and CDC. Existing V2 tables can be upgraded in place with ALTER TABLE without rewriting data.
Using the new types requires a Spark 4.0+ based engine, such as AWS Glue 6.0+ or Amazon EMR 8.1+, and they work across Redshift, EMR and Glue. There is no additional charge. Together with the Aurora integration above, this means you can write variant-typed event data from Spark and read it from Postgres in the same lake.
Read more — AWS News Blog