Cover Image for How are frontier models evaluated?
Cover Image for How are frontier models evaluated?
14 Going
Registration
One Spot Remaining
Hurry up and register before the event fills up!
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

​How do we know what the most advanced AI models are actually capable of?

​Join us for a talk with Martin Milbradt, an AI safety research engineer who works with METR and EquiStamp, evaluating frontier AI models for potentially dangerous capabilities. His work includes designing benchmarks, running evaluations, and improving how we measure what increasingly capable AI systems can do.

​Martin will give us an accessible look into frontier model evaluations: why they matter, how evaluations are designed and conducted, what researchers are trying to measure, and where the current limitations are.

​The talk will be followed by time for questions and discussion.

Location
Möbelkollektiv GmbH
Bauteil C / 3.OG, Wiesentalstraße 40, 90419 Nürnberg, Germany
14 Going