Skip to main content

Cyber Tech Insights

Data Lakehouse Architecture Explained

October 4, 2026
Data Lakehouse Architecture: 5 Best Practices Explained

Sponsored resource. When you request this resource, the details you submit are shared with its sponsor, who may contact you. See our Privacy Policy.

Data lakehouse architecture explained: open table formats on object storage, why organisations adopt it and what to evaluate before you build one.

For years organisations ran two platforms: a data lake for cheap, flexible storage of raw data and a data warehouse for governed, high-performance analytics. The lakehouse pattern aims to deliver both on a single copy of data.

How it works

Data is stored in open file formats such as Parquet on low-cost object storage. An open table format — for example Apache Iceberg, Delta Lake or Apache Hudi — adds a metadata layer that brings database-like capabilities: ACID transactions, schema enforcement and evolution, time travel and efficient updates and deletes. Multiple engines can then query the same tables.

Why organisations adopt it

  • One copy of data for BI, data science and machine learning, reducing duplication.
  • Open formats that reduce lock-in and let you choose query engines.
  • Low storage cost using object storage that scales independently of compute.
  • Support for diverse data, from structured tables to logs and documents.

Typical layers

Many teams organise data into layers: raw data as ingested, cleaned and conformed data, and curated, business-ready data products. Each layer has clear quality expectations and owners.

What to evaluate

  • Which table format and engines fit your skills and existing tools?
  • How will governance, access control and lineage work across engines?
  • What performance do your heaviest BI workloads require?
  • How will you manage costs for compute that scales on demand?
Related: see Object, Block or File Storage? for how object storage underpins this architecture.

5 best practices for building a data lakehouse

  1. Choose an open table format deliberately. Evaluate engine support, community activity, and features such as schema evolution and partition evolution before standardising.
  2. Design layers with clear contracts. Define what data quality, freshness and documentation each layer guarantees, and who owns it.
  3. Centralise governance. Use a catalogue that enforces permissions consistently, whichever engine queries the data, and records lineage.
  4. Optimise file layout. Compact small files, partition sensibly and cluster data on common filter columns to keep queries fast and costs low.
  5. Control compute spend. Separate workloads into right-sized clusters or warehouses, set auto-suspend policies and monitor cost per query or per team.

Lakehouse versus warehouse

A cloud data warehouse remains an excellent choice for structured BI workloads and offers mature performance and management. A lakehouse adds value when you need to combine BI with data science, machine learning and semi-structured data on one open copy, or want to avoid committing all data to a single engine.

Common mistakes to avoid

  • Recreating a data swamp by loading data without owners, documentation or quality checks.
  • Letting each engine apply its own security model, creating inconsistent access.
  • Neglecting table maintenance, which slowly degrades performance.
  • Migrating every warehouse workload at once instead of in measured phases.

Frequently asked questions

Does a lakehouse replace a data warehouse?

Sometimes, but many organisations run both and use the lakehouse for broader data and AI workloads.

Is a lakehouse only for large enterprises?

No. Managed services make the pattern accessible to mid-sized teams, though the governance effort should not be underestimated.

A 90-day action plan

Days 1 to 30: pick one analytics use case with clear value, such as customer churn reporting, and list its sources, consumers and freshness requirements.

Days 31 to 60: land raw sources on object storage in an open table format, build cleaned and curated layers for the use case and connect your catalogue and access policies.

Days 61 to 90: point BI dashboards and a data science notebook at the same curated tables, compare performance and cost with the existing approach, and decide on the next use cases.

Questions to ask platform vendors

  • Which open table formats and query engines are supported without conversion?
  • How are fine-grained permissions enforced across engines?
  • What automatic table maintenance is provided?
  • How is compute priced, and how can we cap spending?
  • How easily can we move our tables to another platform later?

Key terms explained

  • Open table format: a metadata layer such as Apache Iceberg or Delta Lake that adds transactions and schema management to files.
  • Medallion layers: bronze, silver and gold stages that progressively refine information.
  • Time travel: querying a table as it existed at an earlier point.
  • Compaction: merging many small files into fewer large ones for faster queries.

The bottom line

Combining low-cost object storage with open table formats lets organisations serve BI, data science and machine learning from one governed copy of information. The architecture brings flexibility and reduces lock-in, but it still requires disciplined layering, consistent permissions, table maintenance and cost controls. Start with a valuable use case, prove performance and governance, and expand gradually. Many organisations will continue to run warehouses alongside it, choosing the best tool for each workload.

Further reading on data lakehouse

For authoritative, vendor-neutral guidance on data lakehouse, see the Apache Iceberg project. You can also browse our free whitepapers.