Topic
Data Lakehouse
101 posts tagged “Data Lakehouse”.
-
Autonomous Table Optimization When Your Query Workload Stops Being Predictable
Table maintenance used to be a scheduling problem. You knew which tables were hot, you knew that dashboards filtered on order date and region, and…
-
Reading the Apache Iceberg V4 Proposals Before They Land
A Flink job commits every five seconds. Each commit writes one small Parquet file. It also writes a manifest, rewrites a manifest list, and writes a…
-
Building an Honest TCO Model for Open Lakehouses and Proprietary Warehouses
Someone in finance forwards the cloud data warehouse invoice with a one-line question: why is this number growing faster than our data. An architect…
-
How the Iceberg REST Catalog Turned Into the Lakehouse Control Plane
A dbt run updates a fact table and two dimension tables. The fact table commit succeeds. The second dimension commit fails on a conflict. For the…
-
A Migration Playbook for Moving Legacy Warehouses onto Apache Iceberg
The migration plan says twelve weeks. Week fourteen arrives and the team has moved four tables out of six hundred, because table number five turned…
-
Apache Iceberg Support Across the Major Hyperscalers
Every few weeks I get a version of the same question. A team has standardized on Apache Iceberg, they run most of their workloads on one cloud, and…
-
Block vs. Object Storage: A Deep Dive Into the Foundation of Modern Data, and How the Lakehouse Made the Slow Option Fast
Here is one of the strangest and most consequential plot twists in the history of data infrastructure: over the past decade, the analytics industry…
-
Preparing Your Data Lakehouse for the EU AI Act: Auditable Lineage and Data Provenance
The EU AI Act changes the conversation around AI architecture because it makes trust operational. It is not enough to say that an AI system is…
-
Federation and the Lakehouse: Two Roads to Unified Data Access, and How to Know Which One to Take
Every data strategy document written this decade contains some version of the same sentence: we need a single place to access all our data. The…
-
File Encryption for the Lakehouse: The Terminology, the Machinery, and the Hard Problem of Interoperable Encrypted Tables
For years, the open lakehouse had an honest gap that practitioners whispered about and slide decks skipped: encryption. Not the checkbox kind, every…
-
Designing Private, Air-Gapped Data Lakehouses: Scaling Iceberg in Highly Secure, On-Premises Clouds
Some of the most important lakehouse work happens in environments that will never look like a simple public-cloud reference architecture. Defense,…
-
2025 Year in Review Apache Iceberg, Polaris, Parquet, and Arrow
A look back at key developments in Apache Iceberg, Polaris, Parquet, and Arrow in 2025.
-
Comprehensive Hands-on Walk Through of Dremio Cloud Next Gen (Hands-on with Free Trial)
Walkthrough with the new trial of the Dremio Cloud Platform
-
2025-2026 Guide to Learning about Apache Iceberg, Data Lakehouse & Agentic AI
A curated guide to mastering Apache Iceberg, data lakehouse architectures, and the emerging field of Agentic AI for data professionals.
-
An Exploration of the Commercial Iceberg Catalog Ecosystem
Dive into the world of commercial Iceberg catalogs and discover how they enhance data lakehouse architectures for modern data engineering.
-
Building a Universal Lakehouse Catalog - Beyond Iceberg Tables
Exploring paths to a universal lakehouse catalog that supports multiple data formats and engines, building on Apache Iceberg's success.
-
Intro to Apache Iceberg with Apache Polaris and Apache Spark
Learn how to leverage Apache Iceberg with Apache Polaris and Apache Spark to build scalable and efficient data lakehouses.
-
The State of Apache Iceberg v4 - October 2025 Edition
What's Coming in Apache Iceberg v4: A Deep Dive into the Future of Open Table Formats
-
The Ultimate Guide to Open Table Formats - Iceberg, Delta Lake, Hudi, Paimon, and DuckLake
Understanding Iceberg, Delta Lake, Hudi, Paimon, and DuckLake
-
The 2025 & 2026 Ultimate Guide to the Data Lakehouse and the Data Lakehouse Ecosystem
What is the Data Lakehouse and the Data Lakehouse Ecosystem? This comprehensive guide covers everything you need to know about the Data Lakehouse architecture, open table formats like Apache Iceberg, Delta Lake, Apache Hudi, and Apache Paimon, and the modern data ecosystem that supports them.
-
Unlocking the Power of Agentic AI with Apache Iceberg and Dremio
Unlocking the Power of Agentic AI with Apache Iceberg and Dremio
-
The Cost of Neglect – How Apache Iceberg Tables Degrade Without Optimization
Learn how Apache Iceberg tables can degrade over time without optimization and what issues this causes for performance, cost, and governance.
-
How to Discover or Organize Lakehouse & Apache Iceberg Meetups
Guide on How to Be Part of the Lakehouse Community
-
The Data Lakehouse - The Benefits and Enhancing Implementation
Understanding the value of a lakehouse and how to get that value faster
-
2025 Comprehensive Guide to Apache Iceberg
What is Apache Iceberg, How it Works, and Why it Matters!
-
When to use Apache Xtable or Delta Lake Uniform for Data Lakehouse Interoperability
A Guide on when to use Apache Xtable or Delta Lake Uniform for Data Lakehouse Interoperability
-
2025 Guide to Architecting an Iceberg Lakehouse
A Comprehensive Guide to Building a Data Lakehouse with Apache Iceberg
-
10 Future Apache Iceberg Developments to Look forward to in 2025
What is cool about Apache Iceberg's Future
-
Deep Dive into Dremio's File-based Auto Ingestion into Apache Iceberg Tables
Auto ingesting data from JSON, CSV, and Parquet files into Apache Iceberg Tables
-
Dremio, Apache Iceberg and their role in AI-Ready Data
The Role of Dremio and Apache Iceberg in AI-Ready Data
-
Hands-on with Apache Iceberg & Dremio on Your Laptop within 10 Minutes
How to get hands-on with Apache Iceberg
-
Data Modeling - Entities and Events
How to Model Events and Entities
-
All About Parquet Part 01 - An Introduction
All about the Apache Parquet File Format
-
All About Parquet Part 02 - Parquet's Columnar Storage Model
All about the Apache Parquet File Format
-
All About Parquet Part 03 - Parquet File Structure | Pages, Row Groups, and Columns
All about the Apache Parquet File Format
-
All About Parquet Part 04 - Schema Evolution in Parquet
All about the Apache Parquet File Format
-
All About Parquet Part 05 - Compression Techniques in Parquet
All about the Apache Parquet File Format
-
All About Parquet Part 06 - Encoding in Parquet | Optimizing for Storage
All about the Apache Parquet File Format
-
All About Parquet Part 07 - Metadata in Parquet | Improving Data Efficiency
All about the Apache Parquet File Format
-
All About Parquet Part 08 - Reading and Writing Parquet Files in Python
All about the Apache Parquet File Format
-
All About Parquet Part 09 - Parquet in Data Lake Architectures
All about the Apache Parquet File Format
-
All About Parquet Part 10 - Performance Tuning and Best Practices with Parquet
All about the Apache Parquet File Format
-
A Guide to dbt Macros - Purpose, Benefits, and Usage
Learning about dbt Macros
-
Data Lakehouse Roundup 1 - News and Insights on the Lakehouse
What's Going on in the Data Lakehouse Space
-
Getting Started with Data Analytics Using PyArrow in Python
Learning to work with PyArrow to run analytics
-
What is Three-Tier Data (Bronze, Silver, Gold) and How Dremio Simplifies It
Process Data from Raw to Clean Aggregated Data
-
A Brief Guide to the Governance of Apache Iceberg Tables
Controlling Access to your Apache Iceberg Tables
-
Exploring Data Operations with PySpark, Pandas, DuckDB, Polars, and DataFusion in a Python Notebook
Learning to work with Python to ingest and query data
-
Ultimate Directory of Apache Iceberg Resources
Apache Iceberg Education, Tutorials and more!
-
Change Data Capture (CDC) when there is no CDC
Handling Synching Changing Data Across Systems
-
Virtualization + Lakehouse + Mesh = Data At Scale
Combining Centralization and Decentralization for Data at Scale
-
Hands-on with Apache Iceberg on Your Laptop - Deep Dive with Apache Spark, Nessie, Minio, Dremio, Polars and Seaborn
The Evolving Data Lakehouse World
-
Why Data Analysts, Engineers, Architects and Scientists Should Care about Dremio and Apache Iceberg
The Evolving Data Lakehouse World
-
5 Trends in the Data Lakehouse Space
The Evolving Data Lakehouse World
-
Using the alexmerced/datanotebook Docker Image
Setting up a quick and easy data environment for data science and analytics
-
Understanding Apache Iceberg Delete Files
Continuing the Understand Apache Iceberg series, this article delves into the Manifest, a critical component of Apache Iceberg's architecture.
-
Understanding the Apache Iceberg Manifest
Continuing the Understand Apache Iceberg series, this article delves into the Manifest, a critical component of Apache Iceberg's architecture.
-
Understanding the Apache Iceberg Manifest List (Snapshot)
Continuing the Understand Apache Iceberg series, this article delves into the Manifest List, a critical component of Apache Iceberg's architecture.
-
Understanding Apache Iceberg's Metadata.json
The role and content of the metadata.json
-
What Apache Iceberg REST Catalog is and isn't
Understanding Iceberg Catalog Interoperability
-
ACID Guarantees and Apache Iceberg - Turning Any Storage into a Data Warehouse
What are ACID Guarantees? WHy do they matter?
-
Data Lakehouse 101 - The Who, What and Why of Data Lakehouses
The Who, What and Why of Data Lakehouses
-
Understanding the Polaris Iceberg Catalog and Its Architecture
Learn about the new open source Iceberg Catalog in Town
-
Apache Iceberg Reliability
Why Apache Iceberg Works
-
Upcoming Data Talks from Alex Merced (And how to follow)
Come see me talk live at these events
-
Databases Deconstructed - The Value of Data Lakehouses and Table Formats
Building up the Data Lakehouse
-
Video Course - Basics of Lakehouse Engineering - Apache Iceberg, Nessie, Dremio
Introductory Course to Data Engineering for Apache Iceberg Lakehouses
-
Partitioning with Apache Iceberg - A Deep Dive
Benefits of Apache Iceberg Partition Evolution and Hidden Partitioning
-
3 Reasons Data Engineers Should Embrace Apache Iceberg
Benefits of Apache Iceberg
-
Running SQL on your Excel Files From Your Laptop with Dremio
How to run SQL on your Excel files easily
-
A Deep Intro to Apache Iceberg and Resources for Learning More
Learning about Apache Iceberg
-
Understanding the Future of Apache Iceberg Catalogs
Java, Rest and the expanding open lakehouse ecosystem
-
End-to-End Basic Data Engineering Tutorial (Spark, Dremio, Superset)
Ingesting Data and Building BI Dashboards
-
5 Open Source Data Projects You Should Be Following
Apache Iceberg, Apache Arrow, Nessie, Ibis, Substrait
-
5 Reasons Dremio is the Ideal Apache Iceberg Lakehouse Platform
Understanding how catalogs work and which one to choose
-
The Apache Iceberg Lakehouse - The Great Data Equalizer
Disrupting the Snowflake/Databricks status quo
-
10 Reasons to Make Apache Iceberg and Dremio Part of Your Data Lakehouse Strategy
Understanding how catalogs work and which one to choose
-
A deep dive into the concept and world of Apache Iceberg Catalogs
Understanding how catalogs work and which one to choose
-
Introduction to ANSI SQL - Understanding the Syntax and Concepts
Learning the Standard SQL Syntax
-
What is the Data Lakehouse and the Role of Apache Iceberg, Nessie and Dremio?
Understanding the Value of the Data Lakehouse
-
Partitioning Practices in Apache Hive and Apache Iceberg
Deep Dive in Data Lake Table Partitioning
-
Introduction to Data Vault Modeling
Understanding the Data Vault Style of Data Warehouse Modeling
-
Table Format FUD - Thinking Through the Table Format Conversion (Apache Iceberg, Apache Hudi, Delta Lake)
Understanding how to choose a table format
-
Embracing the Future of Data Management - Why Choose Lakehouse, Iceberg, and Dremio?
The Future of Data Platforms
-
Open Lakehouse Engineering/Apache Iceberg Lakehouse Engineering - A Directory of Resources
Resources for learning how to Engineer an Open Data Lakehouse
-
Nessie - An Alternative to Hive & JDBC for Self-Managed Apache Iceberg Catalogs
Nessie is the only open-source catalog implementation specifically for Apache Iceberg.
-
Apache Iceberg, Git-Like Catalog Versioning and Data Lakehouse Management - Pillars of a Robust Data Lakehouse Platform
This is where the combined power of Dremio’s Lakehouse Management features and Project Nessie's catalog-level versioning comes into play.
-
Why Dremio is a must for Apache Iceberg Data Lakehouses
Why is Dremio so useful for Apache Iceberg data lakehouses
-
An In-Depth Overview of Open Lakehouse Tech: Apache Iceberg & Nessie
Organizations are seeking innovative solutions to harness the full potential of their data while maintaining flexibility and avoiding vendor lock-in.…
-
Overview of the Open Lakehouse: Why Dremio?
My cloud infrastructure bill has run wild My datasets have many derived copies for different use cases, which can be complex to maintain and keep…
-
An Approach to Architecting a Lower Cost, Fast and Self-Service Data Lakehouse
There are several goals data architects are perpetually trying to improve upon: Speed: Data Analysts and scientists need data to derive insights to…
-
Handling Cross-Origin Cookies with ExpressJS
Data is becoming the cornerstone of modern businesses. As businesses scale, so does their data, and this leads to the need for efficient data…
-
Creating a Local Data Lakehouse using Spark/Minio/Dremio/Nessie
Data is becoming the cornerstone of modern businesses. As businesses scale, so does their data, and this leads to the need for efficient data…
-
Project Nessie: A Look in the Depths
Once upon a time, in the mystical realm of data lakes, there was a growing problem. The inhabitants of this realm, data scientists, and engineers,…
-
Overview of File Encryption Algorithms for Everyone
Welcome to the thrilling world of file encryption! In this blog post, we'll unravel the secrets of file encryption algorithms and why they are the…
-
Parquet File Compression for Everyone (zstd, brotli, lz4, gzip, snappy)
You know how when you're packing for a trip, you try to stuff as many clothes as you can into your suitcase without breaking the zipper? That's kind…
-
Dremio and Modern Data Architecture: Data Lakes, Data Lakehouses and Data Mesh
Today it can seem like a buzzword onslaught in the data space with terms like Data Mesh, Data Lakehouse, and many more being thrown out with every…
-
What is Nessie and Why as a Data Engineer or Architect you should care?
We need to establish a few things to understand why the open-source data catalog, Project Nessie, matters so much. The amount of data and use cases…
-
Resources for Learning more about Catalog level versioning with Project Nessie & Dremio Arctic (Rollbacks, Branching, Tagging and Multi-Table Txns)
Data Quality, Governance, Observability, and Disaster Recovery are issues that are still trying to discover best practices in the world of the data…
-
5 Reasons Your Data Lakehouse should Embrace Dremio Cloud
How your data lakehouse can expand what's possible with Dremio Cloud.
-
Introduction to The World of Data - (OLTP, OLAP, Data Warehouses, Data Lakes and more)
An accessible high-level guide for data and non-data professionals