Alex Merced's Data, Dev and AI Blog

Topic

data-lake

24 posts tagged “data-lake”.

  1. 7 min read

    Migrating to Apache Iceberg: Strategies for Every Source System

    Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern.

  2. 7 min read

    Hands-On with Apache Iceberg Using Dremio Cloud

    A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics.

  3. 7 min read

    Approaches to Streaming Data into Apache Iceberg Tables

    Stream data into Iceberg with Spark Structured Streaming, Flink, or Kafka Connect. Here is how each works and the trade-offs between latency and maintenance.

  4. 7 min read

    Using Apache Iceberg with Python and MPP Query Engines

    Access Iceberg tables from Python with PyIceberg, DuckDB, and Polars, or through MPP engines like Dremio, Spark, and Trino. Here is how each approach works.

  5. 7 min read

    Apache Iceberg Metadata Tables: Querying the Internals

    Iceberg metadata tables let you query snapshots, files, manifests, and partitions using SQL. Here is every metadata table and how to use them.

  6. 7 min read

    Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup

    Keep Iceberg tables fast with compaction, snapshot expiry, orphan cleanup, and manifest rewriting. Here is when and how to run each operation.

  7. 7 min read

    How Data Lake Table Storage Degrades Over Time

    Iceberg tables degrade through small files, orphan files, metadata bloat, sort order decay, and partition skew. Here is how to diagnose each problem.

  8. 7 min read

    When Catalogs Are Embedded in Storage

    S3 Tables and MinIO AI Stor embed the Iceberg catalog directly in the storage layer. Here is when embedded catalogs make sense and when they do not.

  9. 7 min read

    What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg

    Lakehouse catalogs store metadata pointers, manage namespaces, and enforce access control. Here is the complete catalog landscape from Polaris to Glue.

  10. 8 min read

    Writing to an Apache Iceberg Table: How Commits and ACID Actually Work

    Here is exactly how an engine writes to an Iceberg table, step by step, from data files through the atomic commit that makes ACID guarantees possible.

  11. 8 min read

    Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans

    Iceberg's hidden partitioning separates physical layout from user queries using transform functions.

  12. 8 min read

    Partition Evolution: Change Your Partitioning Without Rewriting Data

    Iceberg lets you change partition schemes without rewriting data. Here is how partition evolution works internally and why Hive-style partitioning could.

  13. 8 min read

    Performance and Apache Iceberg's Metadata

    Iceberg's three-layer metadata tree eliminates directory listing and enables multi-level data skipping. Here is how scan planning actually works.

  14. 8 min read

    The Metadata Structure of Modern Table Formats

    Iceberg uses a metadata tree, Delta Lake uses a transaction log, Hudi uses a timeline. Here is exactly how each format organizes metadata and why it matters.

  15. 9 min read

    What Are Table Formats and Why Were They Needed?

    Table formats like Apache Iceberg solved the ACID, schema, and performance problems that turned data lakes into data swamps. Here is how each one works.

  16. 20 min read

    Introduction to ANSI SQL - Understanding the Syntax and Concepts

    Learning the Standard SQL Syntax

  17. 11 min read

    Introduction to Data Vault Modeling

    Understanding the Data Vault Style of Data Warehouse Modeling

  18. 7 min read

    Table Format FUD - Thinking Through the Table Format Conversion (Apache Iceberg, Apache Hudi, Delta Lake)

    Understanding how to choose a table format

  19. 8 min read

    Embracing the Future of Data Management - Why Choose Lakehouse, Iceberg, and Dremio?

    The Future of Data Platforms

  20. 4 min read

    Open Lakehouse Engineering/Apache Iceberg Lakehouse Engineering - A Directory of Resources

    Resources for learning how to Engineer an Open Data Lakehouse

  21. 5 min read

    Nessie - An Alternative to Hive & JDBC for Self-Managed Apache Iceberg Catalogs

    Nessie is the only open-source catalog implementation specifically for Apache Iceberg.

  22. 6 min read

    Apache Iceberg, Git-Like Catalog Versioning and Data Lakehouse Management - Pillars of a Robust Data Lakehouse Platform

    This is where the combined power of Dremio’s Lakehouse Management features and Project Nessie's catalog-level versioning comes into play.

  23. 7 min read

    5 Reasons Your Data Lakehouse should Embrace Dremio Cloud

    How your data lakehouse can expand what's possible with Dremio Cloud.

  24. 5 min read

    Brief Hands on Intro to Apache Iceberg

    Engineer a Data Lakehouse with Apache Iceberg

Browse all posts

Work with Alex

Menu

Search

Type at least two characters.