Alex Merced's Data, Dev and AI Blog

Topic

streaming

12 posts tagged “streaming”.

  1. 30 min read

    The Hidden Cost of Tiny Iceberg Commits

    Trace what one tiny Iceberg commit writes, then model hourly, per-minute, and per-second cadences so streaming costs become arithmetic, not adjectives.

  2. 31 min read

    Mastering Apache Iceberg v3 Deletion Vectors for High-Throughput Streaming Ingest

    Apache Iceberg v3 deletion vectors for high-throughput streaming ingest: how bitmaps and Puffin files fix CDC write amplification and read decay.

  3. 21 min read

    How Iceberg V3 Deletion Vectors Fixed Merge-on-Read for Streaming Tables

    How Iceberg V3 deletion vectors replaced accumulating positional delete files and made merge-on-read viable for streaming and CDC tables.

  4. 21 min read

    Reading the Apache Iceberg V4 Proposals Before They Land

    A field guide to the Apache Iceberg V4 proposals: adaptive metadata trees, single-file commits, typed statistics, column families, and what is safe.

  5. 31 min read

    Why Iceberg V4 Wants to Retire Equality Deletes, and What Streaming Teams Should Do About It

    Equality deletes made streaming upserts into Iceberg practical at the cost of read performance.

  6. 31 min read

    Apache Fluss and Kafka Solve Different Problems in an Iceberg Pipeline

    Fluss puts a columnar, indexed hot tier between Kafka and Iceberg. Here's what it changes structurally, what Kafka still does better, and how to benchmark.

  7. 30 min read

    Serving Sub-Second Queries Over an Iceberg Lakehouse With a Hot Tier

    A lakehouse cannot serve sub-second queries over seconds-old data. A hot tier in front solves it, with consequences for consistency, governance.

  8. 30 min read

    Freshness Is a Contract, Not a Note on a Dashboard

    Data freshness needs to become an engineering contract with a measurable value, an owner, and consequences.

  9. 30 min read

    The State of Streaming to Apache Iceberg in July 2026: Every Path, Its Latency, and What to Do When Seconds Are Not Fast Enough

    Every path for streaming data into Iceberg in 2026, Flink, Spark, Kafka Connect, broker-native, managed pipelines, with honest latency numbers.

  10. 3 min read

    Optimizing Compaction for Streaming Workloads in Apache Iceberg

    Learn how to design fast, incremental compaction strategies in Apache Iceberg to support high-throughput streaming pipelines without disrupting freshness or performance.

  11. 6 min read

    Data Lakehouse Roundup 1 - News and Insights on the Lakehouse

    What's Going on in the Data Lakehouse Space

  12. 14 min read

    Change Data Capture (CDC) when there is no CDC

    Handling Synching Changing Data Across Systems

Browse all posts

Work with Alex

Menu

Search

Type at least two characters.