

Serving Kimi K3 at scale: An inference deep dive with Together AI
Join Kevin Cui and Austin Silveria from the Together AI inference team for the next session in Inference Hours, moderated by Zain Hasan. This one's about serving Kimi K3 in production.
Our first Kimi K3 session covered what makes the model itself notable: a 3T-class open model with native vision and a 1-million-token context window. This time, we're covering what serving it looks like, starting with why hybrid attention makes caching harder than it looks.
From there, we'll get into the rest of the stack: the quality work, the optimizations, and the performance gains we've seen every week since launch.
This is a systems-focused session for engineers and infra teams, whether you're running K3 on Together or weighing similar tradeoffs for your own stack.
What we'll cover
Architecture: the structural choices that shape how K3 gets served
Quality: the internal tooling that catches issues before they reach production
Optimization: parallelism across hardware, kernel-level fusions, and an in-house speculative decoding model
Performance: throughput and latency gains since launch, week over week
Live Q&A: bring your own questions, or pull from the K3 paper
Register to join the webinar live, ask questions, and receive a recording of the talk.