Alex Merced's Data, Dev and AI Blog

Topic

metadata

10 posts tagged “metadata”.

  1. 30 min read

    Context Engineering for Data Agents

    An executive asks an internal agent what revenue was last quarter. It returns a number, formatted nicely, with the SQL it ran. The number is wrong by…

  2. 30 min read

    Migrating Into Iceberg Without Moving Data

    There are three ways to get an existing dataset into Iceberg and only one of them involves copying files. Teams reach for the copy by default,…

  3. 30 min read

    Partition Statistics Files in Apache Iceberg

    A table with 40,000 partitions takes eleven seconds to plan a query that scans three of them. The scan itself finishes in two. Somebody looks at the…

  4. 29 min read

    Metadata Platforms in 2026: DataHub, OpenMetadata, Atlan, and Catalog Convergence

    The word "catalog" has meant two different things in data infrastructure for about a decade, and in 2026 the two are colliding. The first meaning is…

  5. 31 min read

    Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar

    Picture a table with 40 million data files. Every one of those files has an entry in a manifest, and every entry carries per-column statistics for…

  6. 30 min read

    The Hidden Cost of Tiny Iceberg Commits

    There is a number in your streaming configuration that is quietly deciding your lakehouse's operational future, and it looks completely innocent: the…

  7. 31 min read

    Iceberg v4's Adaptive Metadata Tree, Explained From First Principles

    The best way to understand the centerpiece of the Apache Iceberg v4 design effort is not to read the proposal first. It is to earn the proposal:…

  8. 30 min read

    Why Iceberg v4 Is Really About Making the Cost of Change Proportional to the Change

    Read enough of the Apache Iceberg v4 proposals, the design documents, the dev-list threads, the community sync notes, and a pattern emerges that no…

  9. 21 min read

    Reading the Apache Iceberg V4 Proposals Before They Land

    A Flink job commits every five seconds. Each commit writes one small Parquet file. It also writes a manifest, rewrites a manifest list, and writes a…

  10. 3 min read

    The Endgame – Building an Autonomous Optimization Pipeline for Apache Iceberg

    Learn how to automate compaction, snapshot expiration, and layout optimization in Apache Iceberg using metadata-driven triggers and orchestration tools for a self-healing lakehouse.

Browse all posts

Newsletter

Get new posts in your inbox

Deep dives on Apache Iceberg, lakehouse architecture and applied AI. No spam, unsubscribe anytime.

Subscribe

Menu

Search

Type at least two characters.