Building Iceberg Pipelines in Python Without Standing Up Spark
A data scientist has a transformation that takes forty lines of pandas. It reads two Iceberg tables, joins them, applies a scoring function…
The archive
546 posts on Apache Iceberg, lakehouse architecture, data engineering and applied AI.
A data scientist has a transformation that takes forty lines of pandas. It reads two Iceberg tables, joins them, applies a scoring function…
A platform team I spoke with watched their query volume rise 40 times in six weeks. No new dashboards, no new users, no new data sources.…
A team I talked with recently had hardcoded one model name into forty places in their codebase. The model was deprecated with sixty days…
An inventory agent rerouted a shipment last quarter for a company I spoke with, based on stock levels that were six hours old. The…
Two years ago the interesting question was whether your data platform supported Apache Iceberg. Today every platform claims it does, and…
A colleague spent forty minutes last month provisioning a cluster to profile a 90 GB Parquet dataset. Startup, dependency resolution, a…
Follow an Apache Iceberg design discussion about column-level updates long enough and it stops being an Iceberg discussion. The question of…
A manufacturer in southern Germany asked me a question during an architecture review that I have thought about since. Their data sat in a…
A team I worked with had a dashboard that loaded in three seconds in January and forty seconds in June. Data volume grew 20 percent over…
A finance team asks which of their 40,000 open invoices will pay late. The data sits in a table with 22 columns: customer, terms, amount,…
Newsletter
Deep dives on Apache Iceberg, lakehouse architecture and applied AI. No spam, unsubscribe anytime.