

Optimising Open Source Model Inference
Join Earlybird and Condense for a technical breakfast discussion on the economics of running open source models in production.
There's a common promise that switching to open source cuts your inference bill - but in practice, it often doesn't. Output token prices look cheaper, yet cost per task tends to be higher because open models are more token-intensive. As startups scale, agent economics become one of the most important (and least understood) parts of the stack.
Povilas Nagrockas (CTO and co-founder at Condense) will walk through what they're seeing across teams, with real benchmarks - cost per PR and per work unit, open source vs. frontier models like Fable 5.1 - before opening up to the room. We'll get into where the spend actually goes and how to optimize it: compaction, gateway routing, and model routing among them.
This is a discussion, not a fireside chat. Come ready to share your own examples and push back. Best for engineers and founders working hands-on with LLM infrastructure.
Agenda:
Arrivals: 08:30–09:00
Kickoff & discussion: 09:00–10:00
See past roundtables here:
https://www.akashbajwa.co/t/ai-roundtables
I agree that photographs and/or video may be taken during the event and used by the hosts on social media and in future promotional materials.