Digital Transformation & Artificial Intelligence

Lakehouse Architecture for Unified Analytics and Machine Learning

Equips data engineers and architects to design a lakehouse that serves unified analytics and machine learning from one governed copy of data, with medallion layering.

Duration5 training days
Content4 modules · 8 sessions
On completionAccredited attendance certificate
About the programme

Course Overview

Running a data lake for machine learning and a separate warehouse for reporting means paying twice for storage, twice for pipelines and still arguing about which copy of a number is correct. Lakehouse architecture removes that split by bringing warehouse-grade reliability, through open table formats that add ACID transactions and schema enforcement, directly onto low-cost object storage that both analytics and machine learning workloads can share. This course builds a lakehouse from principles to practice: structuring data through bronze, silver and gold layers so raw, cleaned and business-ready data each have a clear place, managing schema evolution and time travel so historical reports and training datasets stay reproducible, and tuning file layout and clustering so BI query performance holds up as volume grows. Machine learning sessions cover reading features directly from governed tables so training and inference stay consistent. Governance and cost-control sessions close the course, covering cataloguing, access control and the storage tiering decisions that keep a growing lakehouse affordable.

Expected Learning Outcomes

01

Justify a lakehouse architecture against a separate data lake and warehouse for a given workload.

02

Apply an open table format to bring ACID transactions and schema enforcement to object storage.

03

Structure data into bronze, silver and gold layers with clear ownership and transformation rules.

04

Manage schema evolution and use time travel to reproduce a historical report or training dataset.

05

Tune file size, partitioning and clustering to sustain BI query performance as data volume grows.

06

Feed machine learning training and inference from governed lakehouse tables instead of separate extracts.

07

Apply cataloguing, access control and storage tiering to govern cost and risk in a growing lakehouse.

Who Should Attend

01

Data engineers designing or migrating a data platform onto lakehouse architecture.

02

Data architects deciding how to unify currently separate lake and warehouse environments.

03

Machine learning engineers who need consistent, governed features for training and inference.

04

BI and analytics engineers responsible for query performance on growing datasets.

05

Data platform leads accountable for storage and compute cost as data volume increases.

06

Data governance staff extending cataloguing and access control to a new lakehouse environment.

Course Modules

Select any module to see its sessions and points.

01

Principles of Lakehouse Architecture

2 sessions · 8 points

Session 1Combining Data Lake Flexibility With Warehouse Reliability

  • Compare a traditional data lake's flexibility with a data warehouse's transactional reliability to justify the lakehouse approach.
  • Identify workloads currently duplicated across a separate lake and warehouse that a lakehouse could serve from one copy.
  • Assess which existing pipelines can be retired once raw, refined and analytical data share a single storage layer.
  • Plan a phased migration that proves lakehouse reliability on one domain before moving mission-critical reporting.

Session 2Open Table Formats and ACID Transactions on Object Storage

  • Apply an open table format that adds ACID transactions and schema enforcement on top of low-cost object storage.
  • Use table format versioning to support concurrent readers and writers without the corruption risk of a raw file lake.
  • Enforce schema-on-write constraints at the table format layer to stop malformed records reaching downstream consumers.
  • Evaluate table format features such as time travel and rollback against the organisation's audit and recovery needs.
02

Structuring Data Through the Medallion Architecture

2 sessions · 8 points

Session 1Designing Bronze, Silver and Gold Layers for Governed Refinement

  • Land raw source data unchanged in a bronze layer so the original record is always available for reprocessing.
  • Clean, deduplicate and conform data into a silver layer that downstream teams can query with confidence.
  • Curate business-ready gold tables aligned to specific reporting and analytics needs, not raw system structures.
  • Trace a gold-layer figure back through silver to its bronze source when a business user questions a number.

Session 2Managing Schema Evolution and Time Travel Across Layers

  • Manage schema evolution so new source fields can be added without breaking existing queries on a table.
  • Use table format time travel to reproduce a report exactly as it appeared on a previous date.
  • Version transformation logic alongside data so a pipeline change and its data impact stay traceable together.
  • Set a policy for handling breaking schema changes that gives downstream consumers notice before they occur.
03

Serving Unified Analytics and Machine Learning Workloads

2 sessions · 8 points

Session 1Supporting BI Query Performance From a Single Copy of Data

  • Tune file size, partitioning and clustering in gold tables to keep BI query response times acceptable at scale.
  • Apply caching and materialised views for dashboards that repeatedly query the same aggregated gold tables.
  • Benchmark BI query performance on the lakehouse against the legacy warehouse before decommissioning either.
  • Set workload management rules so a heavy ad hoc query cannot degrade performance for scheduled reporting.

Session 2Feeding Machine Learning Pipelines From the Lakehouse

  • Read training data for machine learning directly from silver or gold tables instead of a separate extract.
  • Use table format time travel to reproduce the exact training dataset a deployed model was built on.
  • Feed feature engineering pipelines from the lakehouse so features stay consistent between training and inference.
  • Coordinate data engineers and data scientists on shared table contracts so both workloads read compatible schemas.
04

Governing and Optimising the Lakehouse

2 sessions · 8 points

Session 1Access Control, Cataloguing and Data Governance

  • Apply row- and column-level access control in the lakehouse catalogue so sensitive fields are restricted by role.
  • Register every table in a central catalogue with owner, classification and lineage recorded for audit.
  • Track data lineage from source system through bronze, silver and gold to satisfy regulatory reporting requests.
  • Review access grants periodically so permissions match current roles rather than historical project membership.

Session 2Optimising Storage Layout and Compute Cost at Scale

  • Compact small files produced by frequent writes to keep query planning and storage costs under control.
  • Tier ageing data to lower-cost storage classes while keeping it queryable for infrequent historical analysis.
  • Monitor compute cost by workload to identify a query pattern or job that disproportionately drives spend.
  • Set retention and archiving rules per data domain instead of applying one policy across all lakehouse data.

What the participant receives

4 course modules

A structured syllabus

8 training sessions

across 5 days

32 detailed points

Applied, detailed content

Accredited attendance certificate

On completing the programme

Complete your registration

We will contact you within one business day to confirm.

Ready to start?

Reserve your seat and start building the skill.

Enroll now

Share this course