Avoiding Metadata Bloat with Snapshot Expiration and Rewriting Manifests
Learn how to prevent and clean up metadata bloat in Apache Iceberg by expiring snapshots and rewriting manifests for better performance and manageability.
The archive
546 posts on Apache Iceberg, lakehouse architecture, data engineering and applied AI.
Learn how to prevent and clean up metadata bloat in Apache Iceberg by expiring snapshots and rewriting manifests for better performance and manageability.
Improve query performance in Apache Iceberg by organizing your data layout with sorting and Z-order clustering. Learn how to reduce scan cost and improve filter effectiveness.
Learn how to design fast, incremental compaction strategies in Apache Iceberg to support high-throughput streaming pipelines without disrupting freshness or performance.
Learn how standard compaction works in Apache Iceberg and why bin packing your data files is essential for maintaining query performance and cost efficiency.
Learn how Apache Iceberg tables can degrade over time without optimization and what issues this causes for performance, cost, and governance.
Guide on How to Be Part of the Lakehouse Community
Introduction to the terms in data engineering
Introduction to the terms in data engineering
Introduction to the terms in data engineering
Introduction to the terms in data engineering
Newsletter
Deep dives on Apache Iceberg, lakehouse architecture and applied AI. No spam, unsubscribe anytime.