Topic
data-lake
24 posts tagged “data-lake”.
-
Migrating to Apache Iceberg: Strategies for Every Source System
Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern.
-
Hands-On with Apache Iceberg Using Dremio Cloud
A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics.
-
Approaches to Streaming Data into Apache Iceberg Tables
Stream data into Iceberg with Spark Structured Streaming, Flink, or Kafka Connect. Here is how each works and the trade-offs between latency and maintenance.
-
Using Apache Iceberg with Python and MPP Query Engines
Access Iceberg tables from Python with PyIceberg, DuckDB, and Polars, or through MPP engines like Dremio, Spark, and Trino. Here is how each approach works.
-
Apache Iceberg Metadata Tables: Querying the Internals
Iceberg metadata tables let you query snapshots, files, manifests, and partitions using SQL. Here is every metadata table and how to use them.
-
Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup
Keep Iceberg tables fast with compaction, snapshot expiry, orphan cleanup, and manifest rewriting. Here is when and how to run each operation.
-
How Data Lake Table Storage Degrades Over Time
Iceberg tables degrade through small files, orphan files, metadata bloat, sort order decay, and partition skew. Here is how to diagnose each problem.
-
When Catalogs Are Embedded in Storage
S3 Tables and MinIO AI Stor embed the Iceberg catalog directly in the storage layer. Here is when embedded catalogs make sense and when they do not.
-
What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg
Lakehouse catalogs store metadata pointers, manage namespaces, and enforce access control. Here is the complete catalog landscape from Polaris to Glue.
-
Writing to an Apache Iceberg Table: How Commits and ACID Actually Work
Here is exactly how an engine writes to an Iceberg table, step by step, from data files through the atomic commit that makes ACID guarantees possible.
-
Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans
Iceberg's hidden partitioning separates physical layout from user queries using transform functions.
-
Partition Evolution: Change Your Partitioning Without Rewriting Data
Iceberg lets you change partition schemes without rewriting data. Here is how partition evolution works internally and why Hive-style partitioning could.
-
Performance and Apache Iceberg's Metadata
Iceberg's three-layer metadata tree eliminates directory listing and enables multi-level data skipping. Here is how scan planning actually works.
-
The Metadata Structure of Modern Table Formats
Iceberg uses a metadata tree, Delta Lake uses a transaction log, Hudi uses a timeline. Here is exactly how each format organizes metadata and why it matters.
-
What Are Table Formats and Why Were They Needed?
Table formats like Apache Iceberg solved the ACID, schema, and performance problems that turned data lakes into data swamps. Here is how each one works.
-
Introduction to ANSI SQL - Understanding the Syntax and Concepts
Learning the Standard SQL Syntax
-
Introduction to Data Vault Modeling
Understanding the Data Vault Style of Data Warehouse Modeling
-
Table Format FUD - Thinking Through the Table Format Conversion (Apache Iceberg, Apache Hudi, Delta Lake)
Understanding how to choose a table format
-
Embracing the Future of Data Management - Why Choose Lakehouse, Iceberg, and Dremio?
The Future of Data Platforms
-
Open Lakehouse Engineering/Apache Iceberg Lakehouse Engineering - A Directory of Resources
Resources for learning how to Engineer an Open Data Lakehouse
-
Nessie - An Alternative to Hive & JDBC for Self-Managed Apache Iceberg Catalogs
Nessie is the only open-source catalog implementation specifically for Apache Iceberg.
-
Apache Iceberg, Git-Like Catalog Versioning and Data Lakehouse Management - Pillars of a Robust Data Lakehouse Platform
This is where the combined power of Dremio’s Lakehouse Management features and Project Nessie's catalog-level versioning comes into play.
-
5 Reasons Your Data Lakehouse should Embrace Dremio Cloud
How your data lakehouse can expand what's possible with Dremio Cloud.
-
Brief Hands on Intro to Apache Iceberg
Engineer a Data Lakehouse with Apache Iceberg