Databricks Data Engineering and Pipelines

Build Databricks pipelines with notebooks, PySpark, and Delta Lake - from ADLS Gen2 ingest through medallion layers to Auto Loader, MERGE patterns, and production Jobs.

databricks-data-engineering-and-pipelines

Advanced

Data Engineering

2 Days

Data Engineering

data-engineering

Online
On-site
Hybrid

Databricks Data Engineering and Pipelines

Build Databricks pipelines with notebooks, PySpark, and Delta Lake - from ADLS Gen2 ingest through medallion layers to Auto Loader, MERGE patterns, and production Jobs.

Duration:
2 Days
Rating:
4.8/5.0
Level:
Advanced
1500+ users onboarded

Who will Benefit from this Training?

Training Objectives

Build a high-performing, job-ready tech team.

Personalise your team’s upskilling roadmap and design a befitting, hands-on training program with Uptut

Key training modules

Comprehensive, hands-on modules designed to take you from basics to advanced concepts
Download Curriculum
  • Databricks Notebooks and PySpark for Pipeline Work
    1. Notebook cell types, magic commands (%sql, %python, %md), and widgets for parameterisation
    2. PySpark DataFrames for analysts: transformations versus actions
    3. Common DataFrame operations used in pipeline notebooks
    4. Organising notebook code for reusable Bronze-to-Gold steps
    5. Hands-on: Build and parameterise a pipeline notebook with SQL and Python cells
  • Ingesting Data from ADLS Gen2 into Delta
    1. Reading from ADLS Gen2 with the abfss:// protocol
    2. Loading CSV, JSON, and Parquet into DataFrames
    3. Writing to Delta tables with append, overwrite, and merge modes
    4. Choosing write patterns for raw landing versus curated layers
    5. Validating row counts and table schemas after ingest
  • Medallion Architecture: Bronze, Silver, and Gold
    1. Medallion design principles for lakehouse data engineering
    2. Bronze: land raw ADLS Gen2 data into Delta with minimal transformation
    3. Silver: cleanse, deduplicate, standardise, and cast types
    4. Gold: business aggregations and wide tables ready for reporting and Power BI
    5. Hands-on: Implement a Bronze-to-Silver-to-Gold pipeline on a sample ADLS Gen2 dataset
  • Schema Evolution and Bad-Record Handling
    1. Schema enforcement versus schema evolution: when to allow change
    2. Configuring schema behaviour for evolving source files
    3. Corrupt-record modes: PERMISSIVE, DROPMALFORMED, and FAILFAST
    4. Quarantining bad records without failing the entire pipeline
    5. Practical rules for production schema change management
  • Auto Loader for Incremental Ingestion
    1. Streaming ingestion from ADLS Gen2 into Delta with Auto Loader
    2. Auto Loader versus batch COPY INTO: when to use each
    3. Schema inference and evolution with cloudFiles.schemaLocation
    4. Using the rescue column for unexpected fields
    5. Operational considerations for continuous and triggered streaming jobs
  • Incremental Loads and MERGE INTO for CDC
    1. Full load versus incremental versus Change Data Capture patterns
    2. MERGE INTO for CDC-style upserts in a single statement
    3. Handling inserts, updates, and deletes against Delta targets
    4. Idempotent pipeline design to avoid duplicate Gold metrics
    5. Validating MERGE outcomes with before/after row checks
  • Orchestrating Pipelines with Databricks Jobs
    1. Creating Jobs, adding tasks, and applying cluster policies
    2. Scheduling pipelines for recurring production runs
    3. Multi-task workflows: dependencies, parallel tasks, and conditional branching
    4. Failure handling with retries, on-failure tasks, and email alerts
    5. Passing parameters and task values across notebook jobs
  • Secrets, Identity, and Secure Storage Access
    1. Creating secret scopes and storing credentials safely
    2. Referencing secrets from notebooks and Jobs without hardcoding
    3. Managed identities and service principals for ADLS Gen2 access
    4. Least-privilege patterns for pipeline identities
    5. Checklist for removing secrets from source-controlled notebooks
  • Monitoring, Cost Control, and Production Ops
    1. Using job run history, cluster logs, and event logs for troubleshooting
    2. Job clusters versus all-purpose clusters for pipeline cost efficiency
    3. Sizing clusters and enabling auto-termination for idle workloads
    4. Detecting common failure modes: schema drift, bad files, and permission errors
    5. Operational checklist for handing pipelines to production support
  • Capstone: Scheduled Incremental Pipeline to Gold
    1. Ingest new files from ADLS Gen2 with Auto Loader into Bronze
    2. Apply Silver cleansing and Gold aggregations with MERGE where needed
    3. Orchestrate the flow as a multi-task Databricks Job on a schedule
    4. Secure credentials with secrets or managed identity patterns
    5. Hands-on: Deliver a scheduled incremental pipeline from ADLS Gen2 to a Gold Delta table

Hands-on Experience with Tools

Training Delivery Format

Flexible, comprehensive training designed to fit your schedule and learning preferences
Opt-in Certifications
AWS, Scrum.org, DASA & more
100% Live
on-site/online training
Hands-on
Labs and capstone projects
Lifetime Access
to training material and sessions

How Does Personalised Training Work?

Skill-Gap Assessment

Analysing skill gap and assessing business requirements to craft a unique program

1

Personalisation

Customising curriculum and projects to prepare your team for challenges within your industry

2

Implementation

Supplementing training with consulting support to ensure implementation in real projects

3

Why this course

  • Reliable lakehouse pipelines: Design Bronze-to-Gold flows on Delta that stay consistent from ingest to reporting.
  • Incremental by default: Use Auto Loader and MERGE INTO patterns instead of brittle full reloads.
  • Production orchestration: Schedule multi-task Databricks Jobs with dependencies, retries, and alerts.
  • Secure operations: Manage secrets, identities, monitoring, and cost controls for pipeline workloads.

Training objectives

  • Build data transformation pipelines using Databricks notebooks and PySpark
  • Implement incremental data loading patterns using Delta Lake MERGE INTO and Auto Loader
  • Design and schedule multi-task workflows using Databricks Jobs
  • Apply the medallion architecture (Bronze, Silver, Gold) to structure data engineering pipelines
  • Ingest data from Azure Data Lake Storage Gen2 into Delta tables across pipeline layers
  • Handle schema evolution, bad records, and data quality in pipelines
  • Manage secrets and credentials securely in Databricks pipeline code
  • Monitor, troubleshoot, and maintain production pipelines

Who will benefit

  • Data Engineers
  • Data Analysts
  • Analytics Engineers
  • BI Professionals
  • Cloud Engineers
  • Data Platform Teams

Lead the Digital Landscape with Cutting-Edge Tech and In-House " Techsperts "

Discover the power of digital transformation with train-to-deliver programs from Uptut's experts. Backed by 70,000+ professionals across the world's leading tech innovators.

Frequently Asked Questions

1. What are the pre-requisites for this training?
Faq PlusFaq Minus

The training does not require you to have prior skills or experience. The curriculum covers basics and progresses towards advanced topics.

2. Will my team get any practical experience with this training?
Faq PlusFaq Minus

With our focus on experiential learning, we have made the training as hands-on as possible with assignments, quizzes and capstone projects, and a lab where trainees will learn by doing tasks live.

3. What is your mode of delivery - online or on-site?
Faq PlusFaq Minus

We conduct both online and on-site training sessions. You can choose any according to the convenience of your team.

4. Will trainees get certified?
Faq PlusFaq Minus

Yes, all trainees will get certificates issued by Uptut under the guidance of industry experts.

5. What do we do if we need further support after the training?
Faq PlusFaq Minus

We have an incredible team of mentors that are available for consultations in case your team needs further assistance. Our experienced team of mentors is ready to guide your team and resolve their queries to utilize the training in the best possible way. Just book a consultation to get support.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.