Cover Image for Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips
Cover Image for Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips
Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
96 Went

Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Zoom link: https://us02web.zoom.us/j/82308186562

Talk #0: Introductions and Meetup Updates

by Chris Fregly and Antje Barth

Talk #1: Disaggregated Speculative Decoding with d-Matrix Accelerators by Tom St. John, Head of Applied Research @ Gimlet Labs

Tom will discuss and demonstrate disaggregated speculative decoding using d-Matrix chips on Gimlet Cloud.

Talk #2: Optimizing Inference Engine Performance by Head of Engineering @ Makora

Makora demos how to tune inference engine performance for the latest models and workloads.

Zoom link: https://us02web.zoom.us/j/82308186562

Related Links

Github Repo: cfregly/ai-performance-engineering

O'Reilly Book: https://www.amazon.com/Systems-Performance-Engineering-Optimizing-Algorithms/dp/B0F47689K8/

YouTube: AIPerformanceEngineering

Generative AI Free Course on DeepLearning.ai: https://bit.ly/gllm

Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
96 Went