

Disaggregated Speculative Decoding and Optimizing Inference Engines Across Chips
Zoom link: https://us02web.zoom.us/j/82308186562
Talk #0: Introductions and Meetup Updates
by Chris Fregly and Antje Barth
Talk #1: Disaggregated Speculative Decoding with d-Matrix Accelerators by Tom St. John, Head of Applied Research @ Gimlet Labs
Tom will discuss and demonstrate disaggregated speculative decoding using d-Matrix chips on Gimlet Cloud.
Talk #2: Optimizing Inference Engine Performance by Head of Engineering @ Makora
Makora demos how to tune inference engine performance for the latest models and workloads.
Zoom link: https://us02web.zoom.us/j/82308186562
Related Links
Github Repo: cfregly/ai-performance-engineering
O'Reilly Book: https://www.amazon.com/Systems-Performance-Engineering-Optimizing-Algorithms/dp/B0F47689K8/
YouTube: AIPerformanceEngineering
Generative AI Free Course on DeepLearning.ai: https://bit.ly/gllm