Alex Merced's Data, Dev and AI Blog

Topic

Iceberg

20 posts tagged “Iceberg”.

  1. 16 min read

    REST Catalog Credential Vending for Iceberg

    An in-depth exploration of rest catalog credential vending for iceberg

  2. 18 min read

    REST Catalog V2: Fixing Iceberg Protocol Debt

    An in-depth exploration of rest catalog v2: fixing iceberg protocol debt

  3. 17 min read

    Zero-Copy Mirroring to Open Iceberg Tables

    An in-depth exploration of zero-copy mirroring to open iceberg tables

  4. 15 min read

    Designing Private, Air-Gapped Data Lakehouses: Scaling Iceberg in Highly Secure, On-Premises Clouds

    Some of the most important lakehouse work happens in environments that will never look like a simple public-cloud reference architecture. Defense...

  5. 28 min read

    Lakehouse Table Formats in 2026: Iceberg, Delta Lake, Hudi, Paimon, and DuckLake, How They Work, Where They Stand, and Where They're Going

    The table format war is over, and the table formats are not. Both halves of that sentence are true, both matter, and the tension between them is ex...

  6. 7 min read

    Migrating to Apache Iceberg: Strategies for Every Source System

    Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern.

  7. 7 min read

    Hands-On with Apache Iceberg Using Dremio Cloud

    A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics.

  8. 7 min read

    Approaches to Streaming Data into Apache Iceberg Tables

    Stream data into Iceberg with Spark Structured Streaming, Flink, or Kafka Connect. Here is how each works and the trade-offs between latency and maintenance.

  9. 7 min read

    Using Apache Iceberg with Python and MPP Query Engines

    Access Iceberg tables from Python with PyIceberg, DuckDB, and Polars, or through MPP engines like Dremio, Spark, and Trino. Here is how each approach works.

  10. 7 min read

    Apache Iceberg Metadata Tables: Querying the Internals

    Iceberg metadata tables let you query snapshots, files, manifests, and partitions using SQL. Here is every metadata table and how to use them.

  11. 7 min read

    Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup

    Keep Iceberg tables fast with compaction, snapshot expiry, orphan cleanup, and manifest rewriting. Here is when and how to run each operation.

  12. 7 min read

    How Data Lake Table Storage Degrades Over Time

    Iceberg tables degrade through small files, orphan files, metadata bloat, sort order decay, and partition skew. Here is how to diagnose each problem.

  13. 7 min read

    When Catalogs Are Embedded in Storage

    S3 Tables and MinIO AI Stor embed the Iceberg catalog directly in the storage layer. Here is when embedded catalogs make sense and when they do not.

  14. 7 min read

    What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg

    Lakehouse catalogs store metadata pointers, manage namespaces, and enforce access control. Here is the complete catalog landscape from Polaris to Glue.

  15. 8 min read

    Writing to an Apache Iceberg Table: How Commits and ACID Actually Work

    Here is exactly how an engine writes to an Iceberg table, step by step, from data files through the atomic commit that makes ACID guarantees possible.

  16. 8 min read

    Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans

    Iceberg's hidden partitioning separates physical layout from user queries using transform functions.

  17. 8 min read

    Partition Evolution: Change Your Partitioning Without Rewriting Data

    Iceberg lets you change partition schemes without rewriting data. Here is how partition evolution works internally and why Hive-style partitioning could.

  18. 8 min read

    Performance and Apache Iceberg's Metadata

    Iceberg's three-layer metadata tree eliminates directory listing and enables multi-level data skipping. Here is how scan planning actually works.

  19. 8 min read

    The Metadata Structure of Modern Table Formats

    Iceberg uses a metadata tree, Delta Lake uses a transaction log, Hudi uses a timeline. Here is exactly how each format organizes metadata and why it matters.

  20. 9 min read

    What Are Table Formats and Why Were They Needed?

    Table formats like Apache Iceberg solved the ACID, schema, and performance problems that turned data lakes into data swamps. Here is how each one works.

Browse all posts

Work with Alex

Menu

Search

Type at least two characters.