Cover Image for Tech Talk: How Coupang Leverages Distributed Cache to Accelerate ML Model Training
Cover Image for Tech Talk: How Coupang Leverages Distributed Cache to Accelerate ML Model Training
Avatar for Alluxio
Presented by
Alluxio
Alluxio is the only AI data acceleration platform that works with your existing data and compute infrastructure. Visit www.alluxio.io to learn more.
85 Went

Tech Talk: How Coupang Leverages Distributed Cache to Accelerate ML Model Training

Virtual
Registration
Registration Closed
This event is not currently taking registrations. You may contact the host or subscribe to receive updates.
About Event

PLEASE USE THIS LINK TO REGISTER: https://www.alluxio.io/events/tech-talk-how-coupang-leverages-distributed-cache-to-accelerate-ml-model-training. You will receive a Zoom link shortly after you finish your registration.

🛍️ About Coupang

Coupang is a leading e-commerce company in South Korea, with over 50,000 employees and $20+ billion in annual revenue.

Coupang's AI platform team builds and manages a large-scale AI platform in AWS for machine learning engineers to train models that enhance and customize product search results and product recommendations for its 100+ million customers.

🤖 Tech Talk Overview

As the search and recommendation models evolve, optimizing the underlying infrastructure for AI/ML workloads is essential for the e-commerce business. Coupang's platform team actively sought to improve their model training pipeline to boost machine learning engineers' productivity, publish models to production faster, and reduce operational costs. 

Coupang focused on addressing several key areas: 

  1. Shortening data preparation and model training time

  2. Improving GPU utilization in training clusters in different regions

  3. Reducing S3 API and egress costs incurred from copying large training datasets across regions

  4. Simplifying the operational complexity of storage system management

✨ What You'll Learn

In this tech talk, Hyun Jung Baek, Staff Backend Engineer at Coupang, will share best practices for leveraging distributed cache to power search and recommendation model training infrastructure.

Hyun will discuss:

  • 🏗️ How Coupang builds a world-class large-scale AI platform for machine learning engineers

  • 🔄 How adding distributed caching to their multi-region AI infrastructure improves GPU utilization, accelerates end-to-end training time, and significantly reduces cross-region data transfer costs

  • ☸️ How to simplify platform operations by using Kubernetes and easily deploy the same architecture to new GPU clusters

🧑‍💻 About the Speaker

Hyun Jung Baek is a Staff Backend Engineer at Coupang.


​Join the Alluxio Community

Avatar for Alluxio
Presented by
Alluxio
Alluxio is the only AI data acceleration platform that works with your existing data and compute infrastructure. Visit www.alluxio.io to learn more.
85 Went