

AI by Hand ✍️ Seminar: Frontier - Kimi 3
AI by Hand Seminars by Prof. Tom Yeh
I’ll calculate the math and sketch the architectures by hand ✍️.
1. Gemma 4 (Google): alternating global and local attention, KV-cache, sparse mixture of experts, per layer embedding.
2. Qwen 3.6 (Alibaba, China): long-context scaling techniques and attention optimizations improving practical context length.
👉 3. Kimi 3 (Moonshot AI): Kimi Delta Attention, attention residuals, and extreme MoE sparsity, activating just 16 of 896 experts in a 2.8T-parameter model.
About the Seminar Series
In 2026, I’ve made a personal commitment to teach regular live AI seminar by hand ✍️.
Foundation: intuition, math, and mental models of core AI concepts, for beginners.
Frontier: research papers, algorithms, and architectures used by frontier models, for advanced AI engineers and researchers.
Live attendance is free.