Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live
60-minute walkthrough of how engineering teams build production data pipelines on AWS in a single day. We cover architecture, live demo, and what your team needs to run this independently.
Full lab + screenshots: https://becloudready.com/workshops
Learn how to build an AWS data lake from scratch, live, using Amazon S3, AWS Glue, and Amazon Athena. This is a free, hands-on data engineering workshop, not a slide deck: you'll get sandbox AWS credentials and build a real, working data lake pipeline yourself, step by step, in 90 minutes.
We'll take raw CSV data, catalog it with an AWS Glue Crawler, query it with Amazon Athena SQL, then build a Glue ETL job that transforms it into partitioned Parquet for faster, cheaper queries. You'll see the exact cost difference between querying CSV and Parquet on the same data, measured live.
What you'll learn:
How to build an AWS data lake pipeline from S3 to Athena, hands-on
What an AWS Glue Crawler does and how it catalogs data without moving it
How to write and run a Glue ETL job (PySpark) to convert CSV to Parquet
Why Parquet and partitioning dramatically cut Amazon Athena query costs
The raw → catalog → transform → catalog → query pattern used in real-world AWS data platforms
Who should attend: Data engineers, analytics engineers, cloud engineers, and engineering managers learning AWS data engineering, Glue, Athena, or data lake architecture. No prior Glue/Athena experience needed.
Format: Free, live, 90 minutes, hosted on Microsoft Teams. Sandbox AWS credentials provided, no AWS bill risk. Optional take-home assignment to practice on your own dataset.
Hosted by BeCloudReady (Databricks Registered Partner) and TorontoAI (10,000+ member tech community).