D2X Technologies, Data to Excellence

Enrolling now

AWS Data Engineering

More than just technology. We teach end-to-end solutions.

Trainer
Jaswanth
Duration
60 days
Mode
Live online, small batches
Demo timing
7:30 AM IST

About this course

Learn to design, build and deploy your own data lakehouse on AWS, from data ingestion to deployment. Practical, real-time and project-driven, with GenAI tools like Amazon Q and Kiro IDE to automate pipelines and boost productivity.

AWSPythonPySparkApache IcebergGitLab CI/CDTerraformAmazon QKiro IDE

Who this is for

  • Fresh graduates who want a first job in data engineering
  • ETL, SQL and BI developers moving from on-premise tools to the cloud
  • Software engineers who want to switch into data roles
  • Working professionals who can only study before office hours

What you will be able to do

  • Plan a data architecture and data model for a real business use case
  • Ingest and process data on AWS with Python and PySpark
  • Build a data lakehouse on Apache Iceberg
  • Automate pipelines with GitLab CI/CD
  • Deploy infrastructure to AWS with Terraform
  • Use Amazon Q and Kiro IDE to work faster and solve real problems
  • Monitor, optimise and serve data for dashboards and SQL analytics

Real-time end-to-end project

Build. Deploy. Showcase. Add it to your portfolio.

One complete project in 1 week to 10 days, the same flow production data teams use.

  1. 01

    Ingest

    Real sources

  2. 02

    Process

    PySpark on AWS

  3. 03

    Store

    Apache Iceberg

  4. 04

    Orchestrate

    GitLab CI/CD

  5. 05

    Deploy

    Terraform

  6. 06

    Analyze

    Dashboard / SQL

Key highlights

  • Complete real-time project (1 to 10 days)
  • GenAI tools and Kiro IDE for pipeline automation
  • Industry relevant use cases
  • Best practices, monitoring and optimisation
  • Resume and interview guidance
  • Interview support and work support
  • Lifetime access to materials
  • Doubt support even after the course

Full curriculum

What you will learn, module by module

Every topic below is taught live, with hands-on practice. Tap a module to see its topics.

Modules
10
Topics
53
Duration
60 days
Mode
Online
01Introduction to Data Engineering & Cloud Foundations2 topics

1.1What is Data Engineering?

  • The role and responsibilities of a Data Engineer in a modern organization
  • How Data Engineering fits in the broader data ecosystem (Data Science, Analytics, BI)
  • Course overview: tools covered, learning path, and what to expect each week

1.2Big Data, Tech Stack & Cloud Computing

  • What Big Data is: the 5 Vs and why traditional tools fall short
  • Traditional Data Engineering tech stack: tools, advantages, and limitations
  • Modern Data Engineering tech stack: how the industry has evolved
  • Hadoop ecosystem: HDFS, MapReduce, advantages and disadvantages
  • Apache Spark's role and key advantages in modern data pipelines
  • Cloud Computing fundamentals: IaaS, PaaS, SaaS explained
  • Overview of major cloud providers and why AWS dominates data engineering
02Python for Data Engineering8 topics

2.1Python Setup & Basic Programming (Part 1)

  • Python installation and environment setup
  • IDE setup: VS Code and PyCharm, project creation and interpreter configuration
  • Jupyter Notebook: local setup and Google Colab
  • The print() statement and parameters
  • Variable assignment, comments, and F-strings
  • Arithmetic, conditional, logical, identity, and membership operators
  • Python data types: str, int, float, bool, None

2.2Python Basic Programming (Part 2)

  • Indexing and slicing on strings and sequences
  • String methods: upper, lower, strip, split, replace, and more
  • Introduction to Lists: creation, indexing, and List methods
  • List comprehensions: writing concise, readable data transformations

2.3Lists, For Loops, If/Elif/Else & Dictionaries

  • for loop patterns: iterating over lists, ranges, and dictionaries
  • Conditional logic: if, elif, else decision trees
  • Dictionaries: creation, access, update, delete, and dictionary methods
  • Practical exercises combining loops and conditionals for data processing tasks

2.4Dictionaries Deep Dive & JSON Parsing

  • Advanced dictionary operations: nested dicts, merging, dict comprehensions
  • JSON format: structure and its central role in data engineering
  • Parsing JSON files in Python using the json module
  • Real-world scenario: reading API responses and config files as JSON

2.5Sets, Tuples & Dynamic SQL Query Generation

  • Tuples: immutability, packing/unpacking, and practical use cases
  • Sets: uniqueness, union, intersection, and difference operations
  • Choosing the right data structure: sets vs lists vs tuples vs dicts
  • Generating SQL queries dynamically using Python strings and f-strings
  • Real-world scenario: building parameterised SQL from a configuration dictionary

2.6List Comprehensions & Functions Introduction

  • List comprehensions with conditions and nested loops
  • Why comprehensions are preferred over manual loops for data work
  • Introduction to functions: why modular, reusable code matters
  • Writing user-defined functions with parameters and return values

2.7Functions, UDFs, Lambda Functions, OracleDB & ConfigParser

  • User-Defined Functions (UDFs): parameters, return values, and variable scope
  • Lambda functions: anonymous one-liners with map, filter, reduce
  • Connecting to Oracle Database from Python using cx_Oracle
  • Reading and managing configuration files using configparser
  • Real-world scenario: reading DB credentials from a .ini config file: never hardcode secrets

2.8Python Libraries, File Handling & Error Management

  • Python's built-in standard library: os, sys, datetime, re, csv, json
  • Creating custom libraries and packages: init .py, module structure, imports
  • Real-world scenario: packaging common pipeline utilities as a reusable module
  • File handling: reading and writing Text, CSV, and JSON files
  • Error handling: try, except, raise, finally: building resilient scripts
  • Logging module: log levels, custom loggers, logging vs print in production
  • Automating SQL queries with Python: running queries, fetching results, writing to files
03PySpark for Big Data Processing14 topics

3.1Pandas Basics & Introduction to PySpark

  • Pandas DataFrames: reading, selecting, filtering, and aggregating small datasets
  • Why Pandas does not scale and when to switch to PySpark
  • PySpark introduction: what it is and how it sits within the Spark ecosystem
  • PySpark vs Pandas: performance, scale, and API differences

3.2RDDs, DataFrames & Intro Scenarios

  • Resilient Distributed Datasets (RDDs): the original Spark abstraction
  • Creating RDDs and performing basic transformations (map, filter, flatMap) and actions
  • Text file and log file processing with RDD parallelism
  • Introduction to PySpark DataFrames and why they supersede RDDs for most use cases
  • Practical scenarios: when to use RDDs vs DataFrames

3.3DataFrame Basic Operations on CSV

  • Creating a SparkSession and loading CSV files into DataFrames
  • Schema inference vs manual schema definition with StructType
  • Core DataFrame operations: show, printSchema, count, describe
  • Reading multiple file formats: CSV, JSON, Parquet
  • Writing DataFrames back to files with format and save mode options

3.4Spark Architecture

  • Spark cluster components: Driver Program, Master, Worker Nodes, Executors
  • Cluster Manager types: Standalone, YARN, Kubernetes
  • DAG (Directed Acyclic Graph) Scheduler: how Spark plans execution stages
  • Task Scheduler: distributing tasks to executor slots
  • Transformations vs Actions: the key distinction every Spark developer must know
  • Lazy Evaluation: why Spark waits until an action to execute
  • Catalyst Optimizer and Tungsten Optimizer: how Spark auto-optimises your code

3.5DataFrame Filtering & Introduction to Functions

  • Row filtering with filter() and where(): single and compound conditions
  • Column selection: select, drop, column aliasing with .alias()
  • Adding and modifying columns with withColumn
  • Introduction to PySpark built-in functions: col, lit, cast
  • Null handling: isNull, isNotNull, fillna, dropna

3.6Handling Null Values, Single-Row & Group Functions

  • Strategies for detecting and handling null values in real-world datasets
  • Single-row (scalar) functions: string functions, numeric functions, date functions
  • Group (aggregate) functions: sum, count, avg, min, max
  • groupBy and agg patterns for summarising and reporting data
  • Introduction to Window Functions: what they are and when to use them

3.7Window Functions (Part 1)

  • Window function concepts: partitioning and ordering without collapsing rows
  • Ranking functions: rank(), dense_rank(), row_number()
  • lead() and lag(): accessing the next and previous row values
  • Running totals and cumulative sums using sum() over a window frame
  • Real-world scenarios: top-N per group, deduplication, session analysis

3.8Window Functions (Part 2) & Joins Introduction

  • Advanced window frame specifications: rowsBetween and rangeBetween
  • Percentile and statistical window functions (ntile, percent_rank)
  • Introduction to DataFrame Joins: why joins are central to data engineering
  • Join types: inner, left, right, full outer, left semi, left anti
  • Join conditions: single key, multi-key, and expression-based joins

3.9Joins: Deep Dive & Example Scenarios

  • Practical join scenarios: customer-orders, product lookups, data enrichment
  • Handling duplicate column names after joins
  • Cross joins and when to use them
  • Broadcast joins: forcing small-table broadcast for performance gains
  • Set operators: union, unionByName, intersect, except
  • Real-world scenario: joining transactional data with reference/dimension tables

3.10Spark SQL, Temp Views & Catalog Introduction

  • Creating Temporary Views and Global Temp Views from DataFrames
  • Running SQL queries on DataFrames with spark.sql()
  • When to use Spark SQL vs the DataFrame API
  • Introduction to the Spark Catalog: databases, tables, functions
  • createOrReplaceTempView vs createOrReplaceGlobalTempView: scope and lifetime

3.11Catalog Implementation (Part 1)

  • Spark Catalog API: listing and querying metadata programmatically
  • Creating managed and unmanaged (external) tables in the Spark Catalog
  • Persisting DataFrames as catalog tables for reuse across sessions
  • Partitioned tables: writing and reading partition-pruned data

3.12Catalog Implementation (Part 2)

  • Working with a Hive Metastore-compatible catalog
  • Table properties, table statistics, and catalog introspection
  • How the Spark Catalog connects to the AWS Glue Data Catalog
  • End-to-end scenario: catalog-driven pipeline where schema is read from metadata

3.13Spark Performance Optimization (Part 1)

  • Data Skewness: causes, detection, and fixes (salting, repartitioning)
  • Shuffling in Spark: why it is expensive and how to minimize it
  • Narrow vs Wide Transformations: understanding the execution cost
  • coalesce vs repartition: which to use and when
  • SparkUI deep dive: reading Jobs, Stages, Tasks, and execution plans

3.14Spark Performance Optimization (Part 2)

  • Broadcast Variables and Broadcast Joins in practice
  • Dynamic Resource Allocation: letting Spark scale executors automatically
  • Adaptive Query Execution (AQE): Spark 3.x automatic runtime optimisation
  • Caching and Persistence: cache(), persist(), and storage levels
  • Data format selection: Parquet vs CSV vs JSON vs Avro: performance impact
  • Spark configuration tuning: spark.executor.memory, spark.sql.shuffle.partitions
04Data Lakehouse Architecture on AWS4 topics

4.1Introduction to Data Lakehouse Architecture

  • Data Warehouse: schema-on-write, advantages and disadvantages
  • Data Lake: schema-on-read, flexibility versus governance challenges
  • Data Lakehouse: combining the best of both, why the industry is adopting it

Medallion Architecture

  • Bronze Layer: raw ingestion, no transformations, source of truth
  • Silver Layer: cleansed, validated, and standardized data
  • Gold Layer: business-ready, aggregated, curated data for analytics and reporting
  • Real-world Lakehouse pipeline design patterns and folder structures in S3

4.2AWS Global Infrastructure & Services Overview

  • AWS Global Infrastructure: Regions, Availability Zones, and Edge Locations
  • Why region selection matters for data residency, latency, and compliance
  • EC2: virtual compute | S3: object storage | Lambda: serverless functions
  • EMR: managed Spark clusters | Glue: serverless ETL service
  • Athena: serverless SQL on S3 | Redshift: cloud data warehouse
  • IAM: identity and access control | Secrets Manager: credential management
  • Lake Formation: data lake governance | CloudWatch: monitoring and alerting

4.3AWS Account Setup, IAM Users & User Groups (Part 1)

  • Creating and securing an AWS account: root user best practices and MFA
  • AWS IAM overview: Users, Groups, Roles, and Policies explained
  • Creating IAM Users and User Groups
  • Attaching AWS managed policies to groups
  • Principle of least privilege: granting only required permissions

4.4IAM Roles, Policies & AWS CLI Setup (Part 2)

  • Creating IAM Roles and attaching policies for AWS services (Glue, Lambda, Redshift)
  • Generating Access Keys and Secret Keys for programmatic access
  • Writing Custom IAM Policies in JSON: resource-level permission control
  • AWS CLI installation and configuration: aws configure, named profiles
  • Testing CLI access: verifying identity, listing buckets, running basic commands
05AWS S3: Storage, SDK & Data Ingestion5 topics

5.1S3 Storage Classes, AWS SDK & File Operations

  • S3 Storage Classes: Standard, Intelligent-Tiering, Glacier, Deep Archive
  • Cost implications of each storage class for Bronze / Silver / Gold layers
  • Creating S3 buckets and folder structures aligned to Medallion Architecture
  • Connecting to S3 using AWS CLI: cp, sync, ls, rm commands
  • Connecting to S3 using AWS SDK (Boto3) in Python
  • Uploading, downloading, and listing S3 objects programmatically

5.2Data Ingestion: On-Premises CSV to S3 & S3-to-S3 File Copy

  • Designing an ingestion pattern from on-premises sources to S3 Bronze layer
  • Automating file copy from on-prem to S3 using Python and Boto3
  • S3-to-S3 file copy: moving and archiving objects across buckets and prefixes
  • File naming conventions and date-based partitioning (year=/month=/day=)
  • Logging, error handling, and retry logic for ingestion scripts

5.3Data Ingestion: Oracle Database to S3 & AWS Secrets Manager

  • Extracting data from Oracle Database to S3 using Python and cx_Oracle
  • AWS Secrets Manager: creating and retrieving secrets for database credentials
  • Why hardcoding credentials is a security risk and how Secrets Manager solves it
  • Integrating Secrets Manager into Python data pipeline scripts
  • End-to-end scenario: Oracle extract → write Parquet files to S3 Bronze layer

5.4AWS Lambda Introduction & S3 Event-Driven Ingestion

  • AWS Lambda architecture: functions, triggers, handlers, and execution roles
  • Lambda function structure: handler function, event object, context object
  • S3 Event Notifications: triggering Lambda when a file lands in S3
  • Parsing and processing S3 event records inside a Lambda function
  • Data Lakehouse architecture revision: where Lambda fits in the ingestion flow
  • Real-world scenario: file arrives in S3 Bronze → Lambda validates and moves to Silver

5.5Event-Driven Data Pipelines Using Lambda

  • Building fully event-driven data pipelines with Lambda triggers
  • Chaining Lambda functions: triggering downstream pipeline steps automatically
  • Error handling, dead-letter queues (DLQ), and retry logic in Lambda
  • JSON config-based generic pipeline patterns: one function, many use cases
  • Real-world scenario: Bronze file arrival → Lambda → Silver Glue Job → Gold notification
06AWS Glue, Athena & Apache Iceberg5 topics

6.1AWS Glue: Getting Started

  • AWS Glue architecture: Glue Studio, Data Catalog, Crawlers, Jobs, and Triggers
  • Setting up your first Glue Job in Glue Studio (visual and script mode)
  • Executing PySpark code inside an AWS Glue Job
  • Glue Job parameters: passing runtime arguments to jobs
  • Glue DynamicFrame vs PySpark DataFrame: differences and when to use each
  • Hands-on ETL: read CSV from S3 → transform with PySpark → write Parquet to S3

6.2Trigger Glue Jobs from Lambda & Athena External Tables

  • Invoking a Glue Job programmatically from a Lambda function using Boto3
  • Passing job parameters from Lambda to Glue at runtime
  • Monitoring Glue Job status from Lambda: polling get_job_run
  • AWS Athena introduction: serverless SQL directly over S3 data
  • Glue Data Catalog as the metadata store for Athena queries
  • Creating External Tables in Athena pointing to S3 Parquet files
  • Querying S3 data with standard SQL through Athena: advantages and limitations

6.3Apache Iceberg: Getting Started

  • What Apache Iceberg is and why it was created: open table format for analytics
  • Iceberg vs traditional Hive tables: hidden partitioning and metadata management
  • ACID transactions on the data lake: insert, update, delete, and merge operations
  • Time Travel: querying historical snapshots by timestamp or snapshot ID
  • Schema Evolution: adding, renaming, and dropping columns without table rewrites
  • Integrating Apache Iceberg with AWS Glue and PySpark
  • Creating your first Iceberg table in S3 via Glue and Athena

6.4Raw to Cleansed: Approach & Error Debugging

  • Designing the Raw (Bronze) → Cleansed (Silver) transformation layer architecture
  • Common data quality issues: nulls, duplicates, type mismatches, format errors
  • Building a cleansing framework in PySpark: validation rules and rejection logic
  • Writing clean records to Silver layer and rejected records to a quarantine path
  • Debugging AWS Glue Jobs: reading CloudWatch logs, common PySpark error patterns
  • Real-world scenario: messy CSV from Oracle → cleansing rules → Silver Parquet

6.5Raw to Cleansed: Load Strategies & SCD Type 1

  • Full load vs incremental load strategies: trade-offs and when to use each
  • Implementing a full overwrite (truncate and reload) pattern
  • Implementing SCD Type 1 (upsert / overwrite) using Apache Iceberg MERGE INTO
  • Identifying changed records using hash comparison or updated timestamp columns
  • Handling late-arriving data and out-of-order record scenarios
  • Real-world scenario: customer dimension SCD Type 1 upsert via Iceberg
07End-to-End Incremental Pipeline: Raw → Cleansed → Curated → Redshift6 topics

7.1Incremental Data Pull from Source to Bronze (Raw Layer)

  • Incremental extraction patterns: watermark-based, CDC, and timestamp-based
  • Storing watermarks and last-run state in DynamoDB or S3 control tables
  • Extracting only new and changed records from Oracle using Python and cx_Oracle
  • Writing incremental files to S3 Bronze with date-partitioned folder structure
  • Idempotency design: ensuring pipeline re-runs never produce duplicate data

7.2Bronze to Silver: Incremental Cleansing & Transformation

  • Reading only new partition files from S3 Bronze (partition pruning)
  • Applying data quality rules: null handling, deduplication, type casting, standardisation
  • Merging incremental records into the Silver Iceberg table using MERGE INTO
  • SCD Type 2 implementation: effective_date, end_date, is_current tracking columns
  • Writing cleansed output to S3 Silver layer in Iceberg / Parquet format
  • Updating the watermark after each successful incremental run

7.3Silver to Gold: Curated Layer & Business Aggregations

  • Designing the Gold (Curated) layer for analytical consumption and BI reporting
  • Incremental refresh of Gold layer Iceberg tables using partition overwrite
  • Building aggregation jobs in PySpark: daily, weekly, and monthly summaries
  • Handling late data corrections and reprocessing in the Gold layer
  • Data Modelling for the Gold Layer: Star Schema
  • Fact tables: grain definition, measures, foreign keys to dimensions
  • Dimension tables: descriptive attributes, surrogate keys, natural keys
  • Building the complete Sales model: Orders Fact + Customer, Product, Date Dimensions

7.4Data Warehouse Modelling: Dimensions & Fact Tables

  • Dimensional modelling fundamentals: Star Schema vs Snowflake Schema
  • Designing Fact Tables: grain, additive vs semi-additive measures
  • Designing Dimension Tables: attributes, surrogate keys, natural keys
  • Slowly Changing Dimensions (SCD Type 1 and Type 2) in a Redshift context
  • End-to-end validation: row counts, sum checks, referential integrity tests

7.5AWS Redshift: Architecture & Setup

  • Redshift Data Warehouse vs Glue Data Catalog: complementary roles
  • Provisioned Cluster vs Serverless: architecture, cost, and when to choose each
  • Cluster creation and configuration: node types, VPC, security groups
  • IAM Roles for Redshift: attaching S3 and Glue access
  • Default databases and schemas: information_schema, pg_catalog, public
  • Creating databases, schemas, and tables in Redshift

7.6Loading Gold Layer Data into Redshift & Querying

  • Loading data from S3 Gold layer into Redshift using the COPY command
  • Accessing Glue Catalog (Iceberg) tables from Redshift and loading into Redshift tables
  • Exporting Redshift data back to S3 using the UNLOAD command
  • Running COPY and UNLOAD commands programmatically from PySpark
  • Analytical SQL in Redshift: querying and validating loaded data
  • Incremental load patterns into Redshift: append, merge, and upsert strategies
08Monitoring, Security & Config-Based Pipeline Patterns2 topics

8.1Lambda Deep Dive & Config-Based Generic Pipelines

  • Lambda execution model: cold starts, concurrency limits, memory, and timeout tuning
  • Building JSON config-based generic Lambda functions: one function handles many pipelines
  • Environment variables vs Parameter Store vs Secrets Manager for Lambda configuration
  • Lambda layers: packaging shared dependencies (Boto3, custom utilities)
  • Real-world scenario: a single configurable Lambda that routes to the correct Glue Job

8.2AWS Secrets Manager & CloudWatch Monitoring

  • Secrets Manager: creating, rotating, and retrieving secrets from Python scripts
  • Integrating Secrets Manager with Glue Jobs, Lambda, and PySpark pipelines
  • AWS CloudWatch: log groups, log streams, and metric filters for pipelines
  • Monitoring Glue Job runs, Lambda invocations, and overall pipeline health
  • Setting CloudWatch Alarms for job failures, latency breaches, and data quality issues
  • CloudWatch Dashboards: building a visual pipeline health and SLA overview
09Infrastructure as Code & CI/CD Deployment4 topics

9.1Terraform for AWS Data Infrastructure

  • What Infrastructure as Code (IaC) is and why it matters for data engineers
  • Terraform fundamentals: providers, resources, variables, outputs, and state
  • Setting up Terraform for AWS: configuring the AWS provider and authentication

Resources You Will Provision with Terraform

  • S3 buckets with lifecycle policies for Bronze, Silver, and Gold layers
  • IAM roles and policies for Glue, Lambda, and Redshift
  • AWS Glue Jobs and Glue Catalog databases
  • AWS Lambda functions with S3 event source mappings
  • Redshift Serverless namespace and workgroup
  • CloudWatch log groups and failure alarms
  • terraform init, plan, apply, and destroy: the full IaC workflow
  • Remote state management using S3 backend and DynamoDB state locking

9.2Terraform Modules & Multi-Environment Management

  • Terraform modules: writing reusable, parameterised infrastructure components
  • Creating a data pipeline module that provisions Glue + Lambda + S3 together
  • Managing dev, staging, and prod environments with Terraform workspaces or .tfvars files
  • Terraform variable files: terraform.tfvars and environment-specific overrides
  • Terraform outputs: exposing ARNs and endpoint URLs for use by pipeline scripts

9.3GitLab CI/CD for Data Pipeline Deployment

  • Git fundamentals for data engineers: branching, commits, merge requests, code review
  • GitLab CI/CD pipeline structure: .gitlab-ci.yml, stages, jobs, and runners

CI/CD Pipeline Stages You Will Build

  • Lint & Validate: flake8/pylint for Python, terraform validate for IaC
  • Unit Test: running PySpark transformation tests
  • Terraform Plan: preview infrastructure changes on every merge request
  • Terraform Apply: auto-apply infrastructure on merge to main branch
  • Deploy Glue Jobs: upload PySpark scripts to S3, update Glue Job definitions
  • Deploy Lambda: package and publish Lambda function code
  • Managing AWS credentials securely in GitLab CI/CD using masked environment variables
  • Environment-specific deployments: automatic to dev, manual gate for production
  • Real-world scenario: one merge to main deploys your entire AWS data platform

9.4End-to-End Deployment Walkthrough

  • Step 1: Provision S3 buckets (Bronze / Silver / Gold) and lifecycle policies
  • Step 2: Create IAM roles and policies for all services
  • Step 3: Deploy Glue ETL Jobs (PySpark scripts) and Crawlers
  • Step 4: Deploy Lambda trigger functions and S3 event notifications
  • Step 5: Create Glue Data Catalog databases and Apache Iceberg table definitions
  • Running the full incremental pipeline end-to-end and validating data at each layer
10Databricks, Resume & Interview Preparation3 topics

10.1Databricks

  • Databricks Introduction & UI
  • Databricks CLI setup
  • Introduction to Unity Catalog, Delta Lake
  • Databricks workspace components: Volume, Foreign tables, notebooks, Secret scope, connection, pipeline, jobs etc.
  • Create & orchestrate the pipelines
  • Types of Clusters for Compute

10.2Resume Building for Data Engineering Roles

  • What hiring managers look for in a Data Engineer resume: skills vs experience
  • Structuring your resume: skills section, project descriptions, and impact statements
  • How to present projects from this course as portfolio items
  • Key skills to highlight: PySpark, AWS (Glue / Lambda / Redshift), Python, SQL, Terraform

10.3Interview Preparation & Career Strategy

  • Python coding questions: data structures, OOP, file handling, generators
  • PySpark questions: transformations vs actions, window functions, optimisation
  • SQL questions: window functions, joins, performance tuning, explain plans
  • AWS architecture questions: designing scalable, cost-efficient pipelines
  • System design: design a full data ingestion and Lakehouse pipeline end-to-end
  • Mock interview scenarios and structuring answers using the STAR method
  • Job search strategy, salary negotiation, and career progression in data engineering

Course deliverable

What you will have built

By the end of this course you will have built and deployed a production-grade, end-to-end data pipeline on AWS, from raw source extraction all the way to a queryable data warehouse, with IaC and CI/CD.

Source Extraction
Python, cx_Oracle, Boto3 (S3)
Raw / Bronze Layer
S3, AWS Lambda (event-driven file triggers)
Cleansed / Silver Layer
AWS Glue (PySpark), Apache Iceberg, SCD Type 1 & 2
Curated / Gold Layer
AWS Glue (PySpark), Iceberg, Star Schema modelling
Data Warehouse
AWS Redshift: COPY from S3 / Glue Catalog
Ad-hoc Querying
Amazon Athena: serverless SQL on S3 / Iceberg
Secrets & Config
AWS Secrets Manager, ConfigParser, IAM
Infrastructure as Code
Terraform: S3, Glue, Lambda, Redshift, IAM
CI/CD Deployment
GitLab CI/CD: lint, test, plan, apply, deploy

Projects you will build

Project 1

Real-time end-to-end project

Build a complete data lakehouse in 1 week to 10 days, from real data sources to dashboards. Build it, deploy it, showcase it and add it to your portfolio.

Project 2

Automated pipeline

Every change runs through GitLab CI/CD and deploys to AWS with Terraform, the way production teams work.

Project 3

GenAI assisted build

Use Amazon Q and Kiro IDE to speed up coding, debugging and pipeline automation.

Questions about this course

Do I need prior coding experience?+

Basic SQL helps. Python and PySpark are taught the way data engineers use them, with practice after every session.

What is the real-time project?+

An end-to-end data lakehouse you build over 1 week to 10 days: ingest real data, process it with PySpark on AWS, store it in Apache Iceberg, automate with GitLab CI/CD, deploy with Terraform and analyze with dashboards and SQL.

Will I get access to materials after the course?+

Yes. You get lifetime access to the materials and doubt support even after the course ends.

Is there a free demo before I join?+

Yes. Attend the free live demo to see the full course plan and project flow, and get your questions answered.

Other courses

Same data. Bigger opportunities.

Book the free live demo or WhatsApp us. We will suggest the right path for your background and goal.