Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI
A data engineer changes the merge logic in a pipeline that writes to an Apache Iceberg table. To test it, they run the job against the…
The archive
613 posts on Apache Iceberg, lakehouse architecture, data engineering and applied AI.
A data engineer changes the merge logic in a pipeline that writes to an Apache Iceberg table. To test it, they run the job against the…
The word "catalog" has meant two different things in data infrastructure for about a decade, and in 2026 the two are colliding. The first…
A platform team has 3,000 Iceberg tables in a Hive Metastore and 900 terabytes of Parquet behind them. They are moving to a REST catalog.…
An engineering organization pays its observability vendor by the gigabyte ingested and keeps thirty days of logs because ninety triples the…
The orchestration question used to be simple: Airflow, or something that wanted to be Airflow. It is not simple in 2026. Apache Airflow 3…
For a long time the answer to "we need analytics on our Postgres data" was a pipeline. Replicate the transactional tables into a warehouse…
A Kafka topic carries order events. A producer team adds a field. Downstream, three consumers keep working because the serialization format…
A nightly job joins a 3-billion-row orders table to a 400-million-row customers table on customerid. Both are Apache Iceberg tables. Both…
A team needs to test a new pipeline against a year of production orders. Production has the data. Production also has names, addresses,…
The question gets asked in two registers. One is genuine anxiety from people whose careers are in the balance, and it deserves a straight…
Newsletter
Deep dives on Apache Iceberg, lakehouse architecture and applied AI. No spam, unsubscribe anytime.