Cover Image for Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live
Cover Image for Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live
Community-driven AI & Tech meetups with network of 10K+ members — monthly in Toronto, quarterly in San Francisco & the Bay Area
Hosted By
4 Going

Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live

Virtual
Registration
Welcome! To join the event, please register below.
About Event

60-minute walkthrough of how engineering teams build production data pipelines on AWS in a single day. We cover architecture, live demo, and what your team needs to run this independently.

Full lab + screenshots: https://becloudready.com/workshops

Learn how to build an AWS data lake from scratch, live, using Amazon S3, AWS Glue, and Amazon Athena. This is a free, hands-on data engineering workshop, not a slide deck: you'll get sandbox AWS credentials and build a real, working data lake pipeline yourself, step by step, in 90 minutes.

We'll take raw CSV data, catalog it with an AWS Glue Crawler, query it with Amazon Athena SQL, then build a Glue ETL job that transforms it into partitioned Parquet for faster, cheaper queries. You'll see the exact cost difference between querying CSV and Parquet on the same data, measured live.

What you'll learn:

  • How to build an AWS data lake pipeline from S3 to Athena, hands-on

  • What an AWS Glue Crawler does and how it catalogs data without moving it

  • How to write and run a Glue ETL job (PySpark) to convert CSV to Parquet

  • Why Parquet and partitioning dramatically cut Amazon Athena query costs

  • The raw → catalog → transform → catalog → query pattern used in real-world AWS data platforms

Who should attend: Data engineers, analytics engineers, cloud engineers, and engineering managers learning AWS data engineering, Glue, Athena, or data lake architecture. No prior Glue/Athena experience needed.

Format: Free, live, 90 minutes, hosted on Microsoft Teams. Sandbox AWS credentials provided, no AWS bill risk. Optional take-home assignment to practice on your own dataset.

Hosted by BeCloudReady (Databricks Registered Partner) and TorontoAI (10,000+ member tech community).

https://becloudready.com/workshops/data-engineering

Community-driven AI & Tech meetups with network of 10K+ members — monthly in Toronto, quarterly in San Francisco & the Bay Area
Hosted By
4 Going