

How are frontier models evaluated?
About Event
How do we know what the most advanced AI models are actually capable of?
Join us for a talk with Martin Milbradt, an AI safety research engineer who works with METR and EquiStamp, evaluating frontier AI models for potentially dangerous capabilities. His work includes designing benchmarks, running evaluations, and improving how we measure what increasingly capable AI systems can do.
Martin will give us an accessible look into frontier model evaluations: why they matter, how evaluations are designed and conducted, what researchers are trying to measure, and where the current limitations are.
The talk will be followed by time for questions and discussion.