Cover Image for Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live
Cover Image for Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live
Community-driven AI & Tech meetups with network of 10K+ members — monthly in Toronto, quarterly in San Francisco & the Bay Area
Hosted By
12 Going

Build an AWS Data Lake From Scratch: S3, Glue & Athena, Live

Virtual
Registration
Welcome! To join the event, please register below.
About Event

​60-minute walkthrough of how engineering teams build production data pipelines on AWS in a single day. We cover architecture, live demo, and what your team needs to run this independently.

​Full lab + screenshots: https://becloudready.com/workshops

​Learn how to build an AWS data lake from scratch, live, using Amazon S3, AWS Glue, and Amazon Athena. This is a free, hands-on data engineering workshop, not a slide deck: you'll get sandbox AWS credentials and build a real, working data lake pipeline yourself, step by step, in 90 minutes.

​We'll take raw CSV data, catalog it with an AWS Glue Crawler, query it with Amazon Athena SQL, then build a Glue ETL job that transforms it into partitioned Parquet for faster, cheaper queries. You'll see the exact cost difference between querying CSV and Parquet on the same data, measured live.

​What you'll learn:

  • ​How to build an AWS data lake pipeline from S3 to Athena, hands-on

  • ​What an AWS Glue Crawler does and how it catalogs data without moving it

  • ​How to write and run a Glue ETL job (PySpark) to convert CSV to Parquet

  • ​Why Parquet and partitioning dramatically cut Amazon Athena query costs

  • ​The raw → catalog → transform → catalog → query pattern used in real-world AWS data platforms

​Who should attend: Data engineers, analytics engineers, cloud engineers, and engineering managers learning AWS data engineering, Glue, Athena, or data lake architecture. No prior Glue/Athena experience needed.

​Format: Free, live, 90 minutes, hosted on Microsoft Teams. Sandbox AWS credentials provided, no AWS bill risk. Optional take-home assignment to practice on your own dataset.

​Hosted by BeCloudReady (Databricks Registered Partner) and TorontoAI (10,000+ member tech community).

​https://becloudready.com/workshops/data-engineering

Community-driven AI & Tech meetups with network of 10K+ members — monthly in Toronto, quarterly in San Francisco & the Bay Area
Hosted By
12 Going