Data Engineering on Azure

Build a strong foundation in Azure Data Engineering, from lakehouse architecture to enterprise ingestion and transformation workflows. Learn how to build end-to-end pipelines using ADLS, Data Factory, Databricks, and Synapse while applying security, monitoring, and production-ready operational practices.

data-engineering-on-azure

Intermediate

Data Engineering

5 Days

Data Engineering

data-engineering

Online
On-site
Hybrid

Data Engineering on Azure

Build a strong foundation in Azure Data Engineering, from lakehouse architecture to enterprise ingestion and transformation workflows. Learn how to build end-to-end pipelines using ADLS, Data Factory, Databricks, and Synapse while applying security, monitoring, and production-ready operational practices.

Duration:
5 Days
Rating:
4.8/5.0
Level:
Intermediate
1500+ users onboarded

Who will Benefit from this Training?

Training Objectives

Build a high-performing, job-ready tech team.

Personalise your team’s upskilling roadmap and design a befitting, hands-on training program with Uptut

Key training modules

Comprehensive, hands-on modules designed to take you from basics to advanced concepts
Download Curriculum
  • Azure Data Engineering Architecture Overview
    1. Data engineering lifecycle (ingest → store → transform → serve)
    2. Modern platform patterns (Data Lake, Data Warehouse, Lakehouse)
    3. Batch vs streaming overview
    4. Choosing Azure services per workload
    5. Hands-on: Activity: Design an Azure reference architecture for an analytics use case
  • ADLS Gen2, Data Formats, and Lake Zones
    1. Storage account fundamentals
    2. Containers and folder strategy
    3. Lake zones (raw, cleansed, curated)
    4. Naming and partitioning best practices
    5. CSV vs JSON vs Parquet
    6. Hands-on: Lab/Exercise: Upload datasets + define partition folders and choose partition keys
  • Security Basics: RBAC, Managed Identity, and Key Vault
    1. RBAC basics for storage
    2. Managed Identity introduction
    3. Data access patterns (engineers vs analysts, least privilege)
    4. Key Vault overview
    5. Why secrets should never be hardcoded
    6. Hands-on: Lab: Connect ADF to Key Vault for secret management
  • Azure Data Factory Fundamentals and Copy Activity
    1. What ADF solves (ingestion automation, scheduling, dependency management)
    2. Key components (linked services, datasets, pipelines, triggers)
    3. Activity types (Copy Activity, Validation activity intro, ForEach intro)
    4. Copy from HTTP source
    5. Copy from Blob/ADLS source
    6. Hands-on: Lab: Build ingestion pipeline (source → ADLS raw) + parameterize dataset path with dynamic folders
  • Incremental Loads, Monitoring, and ADF Operations
    1. Full load vs incremental load
    2. Watermark concepts (updated_at, ingestion timestamp)
    3. Using pipeline parameters for incremental runs
    4. Monitor pipeline runs
    5. Activity logs and failure reasons
    6. Hands-on: Lab: Simulate pipeline failure, troubleshoot, and add retry + failure handling strategy
  • Databricks Fundamentals and Spark Transformations
    1. What is Databricks (Spark-based processing, notebooks, jobs)
    2. Clusters and compute basics
    3. Reading from ADLS securely
    4. Data cleaning (null handling, schema casting, deduplication)
    5. Joins and aggregations
    6. Hands-on: Lab: Transform raw orders to cleansed zone + build curated datasets (daily revenue, top customers)
  • Delta Lake, Jobs, and Data Quality Checks
    1. Why Delta in a lakehouse (ACID transactions, schema enforcement, scalable merges)
    2. Delta operations (overwrite vs append)
    3. Merge/upsert concept
    4. Time travel introduction
    5. Convert notebook into job
    6. Hands-on: Lab: Add validation and stop pipeline on quality failure
  • Synapse Fundamentals and Serverless Lake Queries
    1. Query choices on Azure (Databricks queries, Synapse serverless SQL overview)
    2. When you need a warehouse vs lake queries
    3. Synapse components overview (SQL pools, Spark pools concept, pipelines concept)
    4. Serverless SQL vs Dedicated SQL (when to use each)
    5. External tables concept
    6. Hands-on: Lab: Query curated zone + build analytics views (revenue by region, top customers)
  • Warehouse Modeling and ADF + Databricks + Synapse Flow
    1. Star schema basics (fact + dimension)
    2. KPI reporting dataset design
    3. Transformations for warehouse tables
    4. End-to-end architecture (ingest with ADF, transform with Databricks, serve via Synapse)
    5. Automation and reliability approach
    6. Hands-on: Lab: Create orchestrated pipeline concept (ADF triggers Databricks job, curated data queryable in Synapse)
  • Performance, Cost, Security, and Governance
    1. Databricks cluster cost drivers
    2. Synapse query cost awareness
    3. Storage cost optimization basics
    4. RBAC for storage and compute
    5. Managed identities in pipelines
    6. Hands-on: Workshop: Define governance model for data lake access and ownership
  • Observability, Reliability, and Production Practices
    1. Monitoring across services (ADF logs, Databricks job logs, storage metrics)
    2. SLA and freshness tracking
    3. Alerting and incident readiness
    4. Idempotent pipeline design
    5. Backfill strategies
    6. Hands-on: Activity: Build production readiness checklist for Azure pipelines
  • Capstone: End-to-End Azure Data Engineering Pipeline
    1. Build ADLS zones (raw/cleansed/curated)
    2. ADF pipeline for ingestion
    3. Databricks transformation pipeline (Delta output)
    4. Synapse analytics queries/views
    5. Data quality checks + incremental load logic
    6. Hands-on: Capstone deliverables (architecture diagram, notebooks, curated datasets, KPI queries, documentation)

Hands-on Experience with Tools

Training Delivery Format

Flexible, comprehensive training designed to fit your schedule and learning preferences
Opt-in Certifications
AWS, Scrum.org, DASA & more
100% Live
on-site/online training
Hands-on
Labs and capstone projects
Lifetime Access
to training material and sessions

How Does Personalised Training Work?

Skill-Gap Assessment

Analysing skill gap and assessing business requirements to craft a unique program

1

Personalisation

Customising curriculum and projects to prepare your team for challenges within your industry

2

Implementation

Supplementing training with consulting support to ensure implementation in real projects

3

Why this course

  • Enterprise-ready data platforms: Build governed pipelines using Azure-native security and identity controls.
  • Faster integration with Microsoft ecosystem: Strong alignment with Power BI, Synapse, and ADLS.
  • Improved operational resilience: Use managed services for stable scaling and reliability.
  • Better compliance readiness: Azure governance tools support regulated and enterprise workloads.
  • Accelerated analytics delivery: Ship dashboards and insights faster with integrated cloud services.

Training objectives

  • Understand modern Data Engineering architecture on Azure and service selection for workloads.
  • Build a complete data platform using Azure services across ingestion, storage, transformation, orchestration, and serving/analytics.
  • Design and implement a data lake foundation using Azure Data Lake Storage Gen2 (ADLS) with raw/cleansed/curated zones.
  • Build ingestion and orchestration pipelines using Azure Data Factory (ADF) with linked services, datasets, triggers, and monitoring.
  • Transform data using Azure Databricks with Spark fundamentals and Delta Lake patterns.
  • Implement batch and incremental pipelines including full load and watermark-based processing.
  • Build analytics datasets and data warehouse style tables and views for reporting.
  • Apply security and governance practices including RBAC, managed identities, Key Vault integration, and encryption.
  • Implement monitoring and operational practices including pipeline logs, failure handling, and alerting.
  • Deliver an end-to-end capstone project on Azure with validated outputs and documentation.

Who will benefit

  • Data Engineers
  • Analytics Engineers
  • Cloud Engineers supporting data platforms
  • Data Platform Engineers
  • DevOps engineers working with Azure data services
  • BI engineers moving into data engineering

Lead the Digital Landscape with Cutting-Edge Tech and In-House " Techsperts "

Discover the power of digital transformation with train-to-deliver programs from Uptut's experts. Backed by 70,000+ professionals across the world's leading tech innovators.

Frequently Asked Questions

1. What are the pre-requisites for this training?
Faq PlusFaq Minus

The training does not require you to have prior skills or experience. The curriculum covers basics and progresses towards advanced topics.

2. Will my team get any practical experience with this training?
Faq PlusFaq Minus

With our focus on experiential learning, we have made the training as hands-on as possible with assignments, quizzes and capstone projects, and a lab where trainees will learn by doing tasks live.

3. What is your mode of delivery - online or on-site?
Faq PlusFaq Minus

We conduct both online and on-site training sessions. You can choose any according to the convenience of your team.

4. Will trainees get certified?
Faq PlusFaq Minus

Yes, all trainees will get certificates issued by Uptut under the guidance of industry experts.

5. What do we do if we need further support after the training?
Faq PlusFaq Minus

We have an incredible team of mentors that are available for consultations in case your team needs further assistance. Our experienced team of mentors is ready to guide your team and resolve their queries to utilize the training in the best possible way. Just book a consultation to get support.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.