Databricks Lakehouse Fundamentals

Learn Databricks Lakehouse fundamentals — Delta Lake tables, Unity Catalog governance, and ADLS Gen2-backed Bronze-to-Gold analytics patterns for teams new to the platform.

databricks-lakehouse-fundamentals

Intermediate

Data Engineering

2 Days

Data Engineering

data-engineering

Online
On-site
Hybrid

Databricks Lakehouse Fundamentals

Learn Databricks Lakehouse fundamentals — Delta Lake tables, Unity Catalog governance, and ADLS Gen2-backed Bronze-to-Gold analytics patterns for teams new to the platform.

Duration:
2 Days
Rating:
4.8/5.0
Level:
Intermediate
1500+ users onboarded

Who will Benefit from this Training?

Training Objectives

Build a high-performing, job-ready tech team.

Personalise your team’s upskilling roadmap and design a befitting, hands-on training program with Uptut

Key training modules

Comprehensive, hands-on modules designed to take you from basics to advanced concepts
Download Curriculum
  • Lakehouse Architecture vs Warehouse and Data Lake
    1. The evolution of data architectures: data warehouse, data lake, and lakehouse compared
    2. Why the lakehouse exists: lake limitations, ACID needs, and schema enforcement
    3. How open table formats and cloud object storage enable analytics and engineering together
    4. When to choose lakehouse patterns for migration from on-prem warehouses
  • Databricks Workspace, Clusters, Notebooks, and DBFS
    1. Workspace overview: navigation, notebooks, and compute clusters
    2. Interactive versus job-oriented cluster thinking for learning workloads
    3. Databricks File System (DBFS) concepts and common data access patterns
    4. Unity Catalog metastore awareness at the platform level
    5. Organising notebooks for Delta labs and end-to-end exercises
  • ADLS Gen2 as the Lakehouse Storage Layer
    1. How Azure Data Lake Storage Gen2 underpins Databricks lakehouse storage
    2. Landing raw files and referencing paths from Databricks
    3. External storage concepts used later with Unity Catalog external locations
    4. Practical checklist for secure storage access in training environments
  • Delta Lake Foundations: Tables, ACID, and Schema
    1. Delta Lake basics: Parquet files plus the _delta_log transaction log
    2. Schema enforcement and schema evolution behaviour
    3. Creating Delta tables with CREATE TABLE USING DELTA and CTAS
    4. Writing DataFrames to Delta and validating table metadata
    5. Hands-on: Build and query a Delta Lake table from raw ADLS Gen2 data
  • Reading, Writing, Updates, and MERGE on Delta
    1. Reading and querying Delta tables with SQL and PySpark
    2. UPDATE, DELETE, and MERGE INTO for upsert patterns
    3. Append versus overwrite write modes for curated tables
    4. Validating row counts and schemas after DML operations
  • Time Travel, History, RESTORE, OPTIMIZE, and VACUUM
    1. Time travel with VERSION AS OF and TIMESTAMP AS OF
    2. Table history and audit using DESCRIBE HISTORY
    3. Restoring tables to a previous state with RESTORE
    4. OPTIMIZE for compacting small files and improving scans
    5. VACUUM retention, time travel trade-offs, and safe purge practices
  • Z-Ordering and Delta Performance Basics
    1. Z-Ordering introduction: co-locating related data for faster filters
    2. Choosing Z-Order columns for common analytics predicates
    3. Linking OPTIMIZE and Z-Order to query performance improvements
    4. Simple performance checklist for new lakehouse tables
  • Unity Catalog: Namespace, Permissions, RLS, and Masking
    1. Unity Catalog three-level namespace: catalog, schema, and table
    2. Metastore setup and workspace assignment concepts on Azure
    3. GRANT, REVOKE, and SHOW GRANTS on Unity Catalog objects
    4. Row-level security with dynamic views and column masking for sensitive fields
    5. External locations: registering ADLS Gen2 paths for governed external tables
  • Capstone: Bronze–Silver–Gold to a Governed Gold Table
    1. Ingest raw ADLS Gen2 data into Bronze Delta tables
    2. Transform into Silver and business-ready Gold layers
    3. Register and govern the Gold table in Unity Catalog with least-privilege access
    4. Hands-on: Unity Catalog setup, table registration, access control, and end-to-end pipeline walkthrough

Hands-on Experience with Tools

Training Delivery Format

Flexible, comprehensive training designed to fit your schedule and learning preferences
Opt-in Certifications
AWS, Scrum.org, DASA & more
100% Live
on-site/online training
Hands-on
Labs and capstone projects
Lifetime Access
to training material and sessions

How Does Personalised Training Work?

Skill-Gap Assessment

Analysing skill gap and assessing business requirements to craft a unique program

1

Personalisation

Customising curriculum and projects to prepare your team for challenges within your industry

2

Implementation

Supplementing training with consulting support to ensure implementation in real projects

3

Why this course

  • Lakehouse clarity: See how the lakehouse improves on pure data lakes and traditional warehouses.
  • Delta Lake foundation: Create, query, update, and merge reliable tables with ACID guarantees.
  • Governed analytics: Use Unity Catalog for namespace, permissions, row-level security, and masking.
  • Cloud storage ready: Connect Azure Data Lake Storage Gen2 as the storage layer under Databricks.

Training objectives

  • Understand Lakehouse architecture and how it differs from traditional data warehouses and data lakes
  • Navigate the Databricks workspace: clusters, notebooks, DBFS, and the core interface
  • Create, read, update, and merge Delta Lake tables using both SQL and Python
  • Apply Delta Lake advanced features: time travel, table history, RESTORE, OPTIMIZE, and VACUUM
  • Understand Unity Catalog: metastore hierarchy, catalogs, schemas, and tables
  • Apply data governance controls using Unity Catalog: permissions, row-level security, and column masking
  • Connect Databricks to Azure Data Lake Storage Gen2 as an external storage layer
  • Execute an end-to-end analytics use case from raw ingestion through to a governed Delta table

Who will benefit

  • Data Engineers
  • Data Analysts
  • Analytics Engineers
  • BI Professionals
  • Cloud Engineers
  • Data Platform Teams

Lead the Digital Landscape with Cutting-Edge Tech and In-House " Techsperts "

Discover the power of digital transformation with train-to-deliver programs from Uptut's experts. Backed by 70,000+ professionals across the world's leading tech innovators.

Frequently Asked Questions

1. What are the pre-requisites for this training?
Faq PlusFaq Minus

The training does not require you to have prior skills or experience. The curriculum covers basics and progresses towards advanced topics.

2. Will my team get any practical experience with this training?
Faq PlusFaq Minus

With our focus on experiential learning, we have made the training as hands-on as possible with assignments, quizzes and capstone projects, and a lab where trainees will learn by doing tasks live.

3. What is your mode of delivery - online or on-site?
Faq PlusFaq Minus

We conduct both online and on-site training sessions. You can choose any according to the convenience of your team.

4. Will trainees get certified?
Faq PlusFaq Minus

Yes, all trainees will get certificates issued by Uptut under the guidance of industry experts.

5. What do we do if we need further support after the training?
Faq PlusFaq Minus

We have an incredible team of mentors that are available for consultations in case your team needs further assistance. Our experienced team of mentors is ready to guide your team and resolve their queries to utilize the training in the best possible way. Just book a consultation to get support.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.