Cover Image for Serving Kimi K3 at scale: An inference deep dive with Together AI
Cover Image for Serving Kimi K3 at scale: An inference deep dive with Together AI
Avatar for Together AI Calendar
Hosted By

Serving Kimi K3 at scale: An inference deep dive with Together AI

YouTube
Registration
Welcome! To join the event, please register below.
About Event

Join Kevin Cui and Austin Silveria from the Together AI inference team for the next session in Inference Hours, moderated by Zain Hasan. This one's about serving Kimi K3 in production.

Our first Kimi K3 session covered what makes the model itself notable: a 3T-class open model with native vision and a 1-million-token context window. This time, we're covering what serving it looks like, starting with why hybrid attention makes caching harder than it looks.

From there, we'll get into the rest of the stack: the quality work, the optimizations, and the performance gains we've seen every week since launch.

This is a systems-focused session for engineers and infra teams, whether you're running K3 on Together or weighing similar tradeoffs for your own stack.

What we'll cover

  • Architecture: the structural choices that shape how K3 gets served

  • Quality: the internal tooling that catches issues before they reach production

  • Optimization: parallelism across hardware, kernel-level fusions, and an in-house speculative decoding model

  • Performance: throughput and latency gains since launch, week over week

  • Live Q&A: bring your own questions, or pull from the K3 paper

Register to join the webinar live, ask questions, and receive a recording of the talk.

Avatar for Together AI Calendar
Hosted By