Enhancing Python Data Loading in the Cloud for AI/ML
Hybrid event hosted by our friends Big Data Bellevue :)
PLEASE REGISTER IN-PERSON ON MEETUP: https://www.meetup.com/big-data-bellevue-bdb/events/298608625/
Register for live stream: https://us06web.zoom.us/webinar/register/WN_zonHaqaTQvq2fen2jYC_XA
🌟 Calling developers and engineers in Seattle!
🚀 Are you eager to optimize your AI/ML workloads?
Enhancing Python Data Loading in the Cloud for AI/ML
Bin Fan (VP of Open Source & Chief Architect @ Alluxio) will address the critical challenge of optimizing data loading for distributed Python applications within AI/ML workloads in the cloud, focusing on popular frameworks like Ray and Hugging Face. Integration of Alluxio's distributed caching for Python applications is accomplished using the fsspec interface, thus greatly improving data access speeds. This is particularly useful in machine learning workflows, where repeated data reloading across slow, unstable or congested networks can severely affect GPU efficiency and escalate operational costs.
What you will learn:
How to optimize data loading and overall performance of data-intensive Python applications (Ray, HuggingFace, etc.)
How to accelerate data access (data I/O) and maximize GPU utilization
The tech stack of Alluxio+fsspec+Ray/HuggingFace and the benefits of having Alluxio in this stack
Real-world stories optimizing AI/ML workloads in production
Speaker: Bin Fan
Bin Fan is VP of Open Source and Chief Architect at Alluxio. Prior to joining Alluxio as a founding engineer, he worked for Google to build its next-generation storage infrastructure. Bin received his PhD in computer science from Carnegie Mellon University on the design and implementation of distributed systems.