Topic
metadata
10 posts tagged “metadata”.
-
Context Engineering for Data Agents
An executive asks an internal agent what revenue was last quarter. It returns a number, formatted nicely, with the SQL it ran. The number is wrong by…
-
Migrating Into Iceberg Without Moving Data
There are three ways to get an existing dataset into Iceberg and only one of them involves copying files. Teams reach for the copy by default,…
-
Partition Statistics Files in Apache Iceberg
A table with 40,000 partitions takes eleven seconds to plan a query that scans three of them. The scan itself finishes in two. Somebody looks at the…
-
Metadata Platforms in 2026: DataHub, OpenMetadata, Atlan, and Catalog Convergence
The word "catalog" has meant two different things in data infrastructure for about a decade, and in 2026 the two are colliding. The first meaning is…
-
Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar
Picture a table with 40 million data files. Every one of those files has an entry in a manifest, and every entry carries per-column statistics for…
-
The Hidden Cost of Tiny Iceberg Commits
There is a number in your streaming configuration that is quietly deciding your lakehouse's operational future, and it looks completely innocent: the…
-
Iceberg v4's Adaptive Metadata Tree, Explained From First Principles
The best way to understand the centerpiece of the Apache Iceberg v4 design effort is not to read the proposal first. It is to earn the proposal:…
-
Why Iceberg v4 Is Really About Making the Cost of Change Proportional to the Change
Read enough of the Apache Iceberg v4 proposals, the design documents, the dev-list threads, the community sync notes, and a pattern emerges that no…
-
Reading the Apache Iceberg V4 Proposals Before They Land
A Flink job commits every five seconds. Each commit writes one small Parquet file. It also writes a manifest, rewrites a manifest list, and writes a…
-
The Endgame – Building an Autonomous Optimization Pipeline for Apache Iceberg
Learn how to automate compaction, snapshot expiration, and layout optimization in Apache Iceberg using metadata-driven triggers and orchestration tools for a self-healing lakehouse.