Cover Image for Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips
Cover Image for Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips
Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
98 Went

Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

​Zoom link: https://us02web.zoom.us/j/82308186562

​Talk #0: Introductions and Meetup Updates

​by Chris Fregly and Antje Barth

​Talk #1: Disaggregated Speculative Decoding with d-Matrix Accelerators by Tom St. John, Head of Applied Research @ Gimlet Labs

​Tom will discuss and demonstrate disaggregated speculative decoding using d-Matrix chips on Gimlet Cloud.

​Talk #2: Optimizing Inference Engine Performance by Head of Engineering @ Makora

​Makora demos how to tune inference engine performance for the latest models and workloads.

​Zoom link: https://us02web.zoom.us/j/82308186562

​Related Links

​Github Repo: cfregly/ai-performance-engineering

​O'Reilly Book: https://www.amazon.com/Systems-Performance-Engineering-Optimizing-Algorithms/dp/B0F47689K8/

​YouTube: AIPerformanceEngineering

​Generative AI Free Course on DeepLearning.ai: https://bit.ly/gllm

Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
98 Went