Alex Merced's Data, Dev and AI Blog

Topic

Parquet

6 posts tagged “Parquet”.

  1. 30 min read

    Variant Shredding Explained: How Iceberg Gets Columnar Performance From Messy JSON

    The Variant type in Apache Iceberg v3 gets described in one sentence so often that the sentence has started doing damage: "store JSON without a…

  2. 20 min read

    A Migration Playbook for Moving Legacy Warehouses onto Apache Iceberg

    The migration plan says twelve weeks. Week fourteen arrives and the team has moved four tables out of six hundred, because table number five turned…

  3. 30 min read

    The Parquet Versioning Problem, and Why Iceberg Cares About It

    A Spark job writes a table. A Trino query against the same table fails with a decoding error on one column. Nothing in the Iceberg metadata looks…

  4. 27 min read

    A Deep Dive Into File Compression: How Data Gets Smaller, Why Codecs Differ, and What to Actually Use in the Lakehouse

    Somewhere in your data platform right now, a single configuration property is quietly deciding a meaningful percentage of your storage bill, your…

  5. 27 min read

    The File Format Renaissance: Parquet, Lance, Vortex, Nimble, BtrBlocks, and the New Physics of Columnar Storage

    For a decade, the file format layer was the most settled real estate in data. Apache Parquet held the analytical world, ORC held the Hive legacy…

  6. 7 min read

    No Code - Convert XLS/CSV files into Parquet with Dremio

    Convert XLS/CSV Files without having to write python

Browse all posts

Newsletter

Get new posts in your inbox

Deep dives on Apache Iceberg, lakehouse architecture and applied AI. No spam, unsubscribe anytime.

Subscribe

Menu

Search

Type at least two characters.