Azure Data Lake Architecture and Implementation

Design and operate Azure Data Lake Storage Gen2 - zone architecture, RBAC and ACLs, Databricks integration, and lifecycle controls for analytics-ready lake foundations.

azure-data-lake-architecture-and-implementation

Intermediate

Data Engineering

2 Days

Data Engineering

data-engineering

Online
On-site
Hybrid

Azure Data Lake Architecture and Implementation

Design and operate Azure Data Lake Storage Gen2 - zone architecture, RBAC and ACLs, Databricks integration, and lifecycle controls for analytics-ready lake foundations.

Duration:
2 Days
Rating:
4.8/5.0
Level:
Intermediate
1500+ users onboarded

Who will Benefit from this Training?

Training Objectives

Build a high-performing, job-ready tech team.

Personalise your team’s upskilling roadmap and design a befitting, hands-on training program with Uptut

Key training modules

Comprehensive, hands-on modules designed to take you from basics to advanced concepts
Download Curriculum
  • Data Lake Principles and Architecture Choices
    1. What a Data Lake is: schema-on-read, raw data preservation, and scalability
    2. Data Lake versus Data Warehouse versus Lakehouse: when to use each
    3. How ADLS Gen2 fits a migration from an on-premise warehouse
    4. Success criteria for a well-governed analytics lake
  • ADLS Gen2 Hierarchical Namespace and Storage Structure
    1. How ADLS Gen2 extends Blob Storage with a hierarchical namespace
    2. Storage account structure: containers, directories, and files
    3. abfss:// path patterns used by analytics engines
    4. Provisioning checklist for a training or project ADLS Gen2 account
  • Zone Architecture: Raw, Curated, and Consumption
    1. Designing Raw (Bronze), Curated (Silver), and Consumption (Gold) zones
    2. What belongs in each zone and how data should move between them
    3. Separating landing, cleansing, and serving concerns for multi-team lakes
    4. Mapping zones to lakehouse Bronze, Silver, and Gold language
  • Folder Design, Partitioning, and Naming Conventions
    1. Partitioning by date, entity, and source system: best practices and anti-patterns
    2. Folder hierarchy patterns that keep discovery and jobs predictable
    3. Naming conventions and metadata standards for a governed Data Lake
    4. Documenting layout so pipelines and analysts share one mental model
  • File Formats: CSV, JSON, Parquet, and Delta
    1. Trade-offs for storage size, schema, and query performance
    2. When CSV or JSON is acceptable in landing versus curated zones
    3. Why Parquet and Delta dominate analytics and lakehouse workloads
    4. Choosing formats per zone without locking future consumers out
    5. Hands-on: Provision an ADLS Gen2 account and design a zone-based folder structure
  • Identity, RBAC, and ACLs on ADLS Gen2
    1. Entra ID (Azure AD) integration for identity and access
    2. RBAC built-in roles: Storage Blob Data Reader, Contributor, Owner, and when to use each
    3. ACLs at directory and file level, and how they differ from RBAC
    4. Combining RBAC and ACLs for enterprise Data Lake access patterns
  • Service Principals, Managed Identities, and Databricks Access
    1. Service principals and managed identities for pipelines and Databricks clusters
    2. Mounting or accessing ADLS Gen2 with service principal auth and credential passthrough
    3. Reading and writing from Databricks using the abfss:// protocol
    4. Unity Catalog external locations for governed ADLS Gen2 paths
  • Lifecycle Management, Retention, and Monitoring
    1. Access tier policies: Hot, Cool, Archive, and automated tiering rules
    2. Retention policies, soft delete, and versioning for data protection
    3. Monitoring with Azure Monitor, Storage Insights, and diagnostic logging
    4. Cost and operations checklist for ongoing lake ownership
  • Capstone: Secure Zone-Based Lake Connected to Databricks
    1. Finalise zone layout, partitions, and naming for a sample analytics domain
    2. Apply RBAC and ACL patterns for engineers versus analysts
    3. Connect Databricks, read partitioned data, and validate access boundaries
    4. Hands-on: Configure RBAC and ACLs, connect Databricks to ADLS Gen2, and read partitioned data

Hands-on Experience with Tools

Training Delivery Format

Flexible, comprehensive training designed to fit your schedule and learning preferences
Opt-in Certifications
AWS, Scrum.org, DASA & more
100% Live
on-site/online training
Hands-on
Labs and capstone projects
Lifetime Access
to training material and sessions

How Does Personalised Training Work?

Skill-Gap Assessment

Analysing skill gap and assessing business requirements to craft a unique program

1

Personalisation

Customising curriculum and projects to prepare your team for challenges within your industry

2

Implementation

Supplementing training with consulting support to ensure implementation in real projects

3

Why this course

  • Architecture that scales: Design ADLS Gen2 lakes with zones, partitions, and naming that stay maintainable.
  • Secure by design: Combine RBAC, ACLs, service principals, and managed identities for enterprise access.
  • Analytics-ready storage: Connect Databricks with abfss:// and Unity Catalog external locations.
  • Cost and lifecycle control: Apply tiering, retention, versioning, and monitoring for production lakes.

Training objectives

  • Understand Data Lake architecture principles and how ADLS Gen2 implements them
  • Design a well-structured Data Lake using zones, folder hierarchies, and naming conventions
  • Configure access control using both RBAC and ACLs on ADLS Gen2
  • Integrate Azure Data Lake Storage Gen2 with Azure Databricks for analytics workloads
  • Implement data lifecycle management: tiering, retention policies, and cost optimisation
  • Apply data governance and security best practices on a Data Lake
  • Register ADLS Gen2 paths as governed storage using Unity Catalog external locations
  • Understand the role of the Data Lake in a migration from an on-premise data warehouse

Who will benefit

  • Data Engineers
  • Data Analysts
  • Analytics Engineers
  • BI Professionals
  • Cloud Engineers
  • Data Platform Teams

Lead the Digital Landscape with Cutting-Edge Tech and In-House " Techsperts "

Discover the power of digital transformation with train-to-deliver programs from Uptut's experts. Backed by 70,000+ professionals across the world's leading tech innovators.

Frequently Asked Questions

1. What are the pre-requisites for this training?
Faq PlusFaq Minus

The training does not require you to have prior skills or experience. The curriculum covers basics and progresses towards advanced topics.

2. Will my team get any practical experience with this training?
Faq PlusFaq Minus

With our focus on experiential learning, we have made the training as hands-on as possible with assignments, quizzes and capstone projects, and a lab where trainees will learn by doing tasks live.

3. What is your mode of delivery - online or on-site?
Faq PlusFaq Minus

We conduct both online and on-site training sessions. You can choose any according to the convenience of your team.

4. Will trainees get certified?
Faq PlusFaq Minus

Yes, all trainees will get certificates issued by Uptut under the guidance of industry experts.

5. What do we do if we need further support after the training?
Faq PlusFaq Minus

We have an incredible team of mentors that are available for consultations in case your team needs further assistance. Our experienced team of mentors is ready to guide your team and resolve their queries to utilize the training in the best possible way. Just book a consultation to get support.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.