

Frontier AI Speaking Club: Evals, Benchmarks vs. Real Use Cases
Frontier AI models are getting better — but how do we know whether benchmarks actually measure what matters in real-world use?
At this edition of Frontier AI Speaking Club, we’ll explore the gap between Evals & Benchmarks and real-world AI applications.
The discussion
Benchmarks are useful for comparing models, but real applications often involve very different challenges: Long, messy, domain-specific documents Retrieval and context that standard benchmarks don't capture Multi-step agentic workflows Evaluation criteria that are difficult to specify in advance
The difference between answering a benchmark question and actually completing a useful task
Case study: IPO Finance Agent (arXiv:2606.23032)
We'll use IPO Finance Agent as a concrete example. The project extends the Finance Agent v2 benchmark toward IPO due diligence, using the SpaceX S-1 filing as a real-world test case.
The broader question: What happens when we stop optimizing for benchmark performance and start evaluating whether an AI system can actually do useful work?
From benchmarks to real-world systems
We'll discuss:
Evals — What should we actually measure? Benchmarks — When are standardized benchmarks useful, and where do they break down?
Real use cases — How should we evaluate agents operating in realistic environments?
Retrieval & context — How much of an evaluation is really testing the model versus the surrounding system?
Automated evaluation — Can LLM-generated rubrics reliably evaluate increasingly complex agentic tasks?
Open projects — What new benchmarks, agents, and evaluation frameworks could meaningfully bridge the gap?
Token API sponsorships
There are also sponsorship opportunities for the most promising projects working on this problem. Projects developing new benchmarks, eval frameworks, agents, datasets, or real-world AI applications may be eligible for token API sponsorships from Z AI to help them run larger-scale experiments and evaluations
Come with a project, an idea, or simply curiosity about where frontier-model evaluation is heading.
Who should come?
AI researchers · engineers · founders · agents builders · eval/benchmark researchers · students · anyone building with frontier models
Format: informal speaking club + discussion + networking
📍 Bubble café, Tereshchenkivska 8, Kyiv — inside Shevchenko Park
🗓 Monday, September 28, 2026
🕕 18:00