In a significant advancement for cloud-native data management, Amazon S3 Tables has announced full support for all data types and capabilities within the Apache Iceberg V3 specification. This update allows organizations to transition from legacy V2 table formats to the more robust V3 standard, unlocking advanced features such as deletion vectors, row-level lineage, and native support for complex data types including variant, nanosecond timestamps, geometry, and geography. By integrating these capabilities directly into S3 Tables, AWS is providing a streamlined path for data engineers to enhance query performance, reduce storage costs, and simplify the management of petabyte-scale analytics datasets.
Apache Iceberg has solidified its position as the industry standard for managing massive analytical datasets in data lakes. Its ability to provide SQL-like functionality—such as schema evolution, hidden partitioning, and time-travel queries—while maintaining data in open-source Parquet files has made it a cornerstone of modern data architecture. Amazon S3 Tables is purpose-built to complement this ecosystem, offering a managed storage layer that handles the heavy lifting of table maintenance, such as automatic compaction, data replication, and Intelligent-Tiering. However, as datasets grow into the billions of rows, the limitations of the previous V2 specification have increasingly become a bottleneck for high-performance analytics teams.
The move to Iceberg V3 addresses several long-standing operational pain points that teams frequently encounter when scaling their data lakes. Under the older V2 specification, performing routine data deletions—such as removing 50,000 user records from a massive 2-billion-row table—could result in the creation of numerous positional delete files. These files, while functional, impose a performance tax on subsequent queries, as the engine must process these files during read operations until a manual or scheduled compaction cycle is completed. Similarly, semi-structured data, such as JSON-based event logs, often required cumbersome workarounds like encoding data into strings, forcing every query to perform costly parsing operations. Other data types, such as geospatial coordinates or high-precision timestamps, were often relegated to suboptimal integer or string formats, further driving up storage costs and query latency.
The V3 specification fundamentally changes this landscape by introducing native support for these complex data types and more efficient row-level operations. One of the most impactful additions is the introduction of deletion vectors. Unlike the positional delete files of the V2 era, which could proliferate and degrade performance, deletion vectors utilize a compact binary format. When a compliance-related deletion occurs, the system writes a single, efficient deletion vector file rather than thousands of individual small files. This drastically reduces the overhead for compaction processes and ensures that query engines can maintain high performance even as the table undergoes frequent modifications.
Complementing this efficiency is the introduction of built-in row lineage, which assigns a unique _row_id and _last_updated_sequence_number to every record. This allows data pipelines to identify and process only the rows that have changed since a specific point in time, effectively eliminating the need for full table scans during incremental data processing. By enabling this surgical approach to data updates and retrieval, organizations can significantly lower their compute consumption and accelerate their data refresh cycles.

For teams dealing with semi-structured data, the new variant data type represents a major leap forward. By storing semi-structured information in a columnar format, the system can automatically shred the data into hidden columns and collect statistics during the write process. When a query is executed, the engine utilizes these internal statistics to perform intelligent file pruning, which substantially reduces I/O compared to the traditional method of parsing raw JSON strings at read time. This native handling of diverse data shapes—such as retail event streams that include varying structures for purchases, page views, and searches—allows developers to store all event types in a single, unified table without the constant overhead of managing complex schema evolution.
Transitioning to these new capabilities is designed to be as seamless as possible. Users can either create new V3 tables from scratch or perform an in-place upgrade of existing V2 tables. This atomic upgrade process does not require the underlying data to be rewritten, minimizing disruption to ongoing analytics workloads. AWS has ensured that the migration path is robust, allowing existing V2 readers to continue interacting with the table during the transition phase until the organization is prepared to fully leverage the V3 feature set. It is important to note, however, that while upgrading from V2 to V3 is a straightforward, one-way operation, the Apache Iceberg specification does not currently support downgrading from V3 to V2. As such, teams are encouraged to verify that all analytical engines accessing the table are compatible with the V3 specification before initiating the upgrade.
The integration of Iceberg V3 into the broader AWS ecosystem highlights the company’s commitment to providing a cohesive data analytics stack. Because AWS supports the Iceberg REST Catalog (IRC) API, these tables maintain interoperability across a wide range of services. Data ingested and stored in Amazon S3 Tables can be processed using Amazon EMR Spark, managed through the AWS Glue Data Catalog, and analyzed via Amazon Redshift. This multi-layered support ensures that organizations can build end-to-end pipelines that benefit from the performance gains of V3 without being locked into a single compute engine.
For developers and architects looking to integrate these features, the practical implementation is straightforward. Configuring a table for the V3 format involves setting specific table properties, such as defining the format version and configuring the write mode to utilize the new deletion vector capabilities. Once these properties are set, the system takes over the routine maintenance, automatically handling the compaction of deletion vectors and ensuring that row lineage fields remain accurate. This automation allows data engineers to shift their focus from infrastructure maintenance to higher-value tasks, such as optimizing query performance and building sophisticated data models that were previously too complex to maintain at scale.
The availability of Iceberg V3 support across all AWS regions where S3 Tables are currently supported marks a significant milestone in the evolution of cloud data storage. By addressing the technical debt inherent in previous table formats, Amazon is enabling organizations to handle ever-increasing volumes of data with greater efficiency and lower operational complexity. As enterprises continue to rely on open standards to avoid vendor lock-in, the robust implementation of the Iceberg specification within the S3 ecosystem provides a reliable, high-performance foundation for the next generation of data-driven applications. Standard S3 Tables pricing remains in effect for these new features, ensuring that the performance and governance benefits of the V3 specification are accessible to all users of the platform.

