

AAIF Berlin · Small Language Models Night
Agents at scale keep sending the same few steps to a frontier model thousands of times a day, and the latency and the bill climb with every call. Swap in a cheap off-the-shelf model and quality drops. Move the model on-device and you run into runtimes and hardware that each behave differently.
This AAIF Berlin night asks one question: what does it take to replace a frontier-model call with a small model you build and run yourself?
We look at finding repeatable steps in production traces, synthetic data and evals for post-training, on-device runtimes like LiteRT-LM, MediaPipe and ONNX, and how CPU, GPU and NPU differ in practice.
The room is small and technical. Come if you build, ship, or pay the inference bill for agentic or LLM systems, or if you want to see how other people do it.
What We'll Cover
Jacek opens with how to post-train a small language model from your own production traces, ending with a live build. Sasha then covers the on-device AI stack, from choosing a model to runtimes and hardware. Each talk is followed by Q&A.
Agenda
18:00 — Doors open, introductions & networking
18:30 — Talk 1: Jacek Gołębiowski (45 min + 15 min Q&A)
19:30 — Talk 2: Sasha Denisov (45 min + 15 min Q&A)
20:30 — Networking
21:00 — Close
Speakers
Jacek Gołębiowski — distil labs. In "Post-trained small language models, and why you'd need one", Jacek shows how to handle the agent steps that are too slow and expensive on a frontier model but not good enough on the cheap tier: finding the step in production traces, generating synthetic data, training and evals, and running the resulting model reliably in production. The talk closes with a live case, building a model in a few hours that scouts thousands of tech posts a day on X. (45 min)
Sasha Denisov. In "Everything You Need to Ship an Edge AI Feature", Sasha gives a layer-by-layer tour of the on-device AI stack: which small models are worth using, how LiteRT-LM, MediaPipe and ONNX differ, and where CPU, GPU and NPU actually diverge, along with open-source tools for tuning, converting and running a model across platforms. (45 min)
Who Should Come
Engineers who build or operate agent and LLM systems, and mobile and edge developers bringing AI on-device, should attend. Researchers and technical leaders weighing inference cost and latency are also welcome. No prior knowledge is required.
Thanks to distil labs for hosting us.
About AAIF
The Agentic AI Foundation (AAIF) is a community dedicated to advancing the understanding and practical application of agentic AI. We bring together engineers, researchers, founders, builders, and AI enthusiasts to explore how autonomous AI systems are designed, evaluated, and deployed.
Through reading groups, workshops, and community discussions, AAIF brings together engineers, researchers, founders, and AI enthusiasts to exchange ideas, challenge assumptions, and learn from one another.
By attending you agree to our Code of Conduct and Privacy Policy.