

Build production-grade AI apps with Nebius Token Factory
Hosted by: Draper Startup House, Nebius Startup Program and Basecamp by Elevation!
This session is a technical walkthrough of what it takes to run Open Frontier AI in production: inference, SFT, custom speculative decoding, and dedicated endpoints, on Nebius Token Factory.
What to expect
A technical presentation and demo of how to run full-stack AI Engineering Pipeline with Nebius Token Factory.
1. Choosing your model stack — open weights vs. proprietary APIs
A decision framework, not a vendor pitch. When open-weight models win on cost, latency, data control, or flexibility, and when they don't. What Kimi K3 changes about that calculus now that a 2.8T MoE model ships at open weights and scores 57 on Artificial Analysis' Intelligence Index, just two points behind GPT-5.6 Sol.
2. Building on Kimi K3
The architecture that makes it work: 2.8T total parameters activating only 16 of 896 experts per token, Kimi Delta Attention, native vision, a full 1M-token window. Where K3 outperforms alternatives on long-horizon coding and agentic tasks, where it doesn't, and integration patterns through Token Factory's OpenAI-compatible API.
3. Full-stack AI engineering — demoed across the whole pipeline
Inference — Fast vs. Base flavors, time-to-first-token under load, multi-region routing, and when to pay for which
Post-training (SFT) — LoRA and full fine-tunes on your own data, and promoting a tuned model straight into production
Custom speculative decoding — training your own draft model against your workload's actual traffic shape, and what that buys you on latency
Dedicated endpoints — performance isolation, reserved capacity, autoscaling, and the point in your growth curve where the switch pays for itself
4. Open floor
Talk to us and your peers about the hardest AI Problems you are trying to solve.
Who this is for
CTOs, VPs of Engineering, AI Architects, AI FDEs and senior AI engineers at growth-stage startups building AI-native or AI-augmented products — typically Series A and beyond.
If you're running inference in production, or about to be, you'll get the most out of this. If you're pre-prototype, this will move faster than is useful.
Room capped at 50 so the discussion stays specific.
What you'll leave with
A decision framework for open-weight vs. proprietary models, mapped to product lifecycle stage
A clear read on Kimi K3's capabilities and where it fits in your model stack
Production-tested patterns for reliable, cost-efficient inference
A concrete view of how post-training and custom speculation change your latency and cost numbers
Code samples and cookbooks from everything demoed, so your team can reproduce it on Monday
[Access to Nebius Builders Program for Token Factory credits to get started with the frontier Open AI Models]
A handful of peers solving the same scaling problems, in the same city, in the same week
Who's hosting
The Nebius Token Factory team — Nebius' production platform for running, fine-tuning, and serving open-source models, and Kimi's Day 0 launch partner for K3.
In partnership with Draper Startup House, Nebius Startup Program and Elevation Capital's Basecamp.
Logistics
Date: 07-Aug-2026 (Friday)
Time: 730PM-10PM (Dinner is served!)
Nothing to bring. Slides, code samples, and cookbooks go out to attendees afterwards. Drinks and Snacks provided.
This event is part of Basecamp, a community-powered tech week in Bengaluru bringing founders, builders, and operators together. Basecamp is free for the community and powered by the community.