PROBE Evaluation Design Jam
The Gap Between Capability and Verified Safety is Widening
Agentic AI systems are shipping faster than we can evaluate them. Meanwhile, regulators across the EU, US, Singapore, and China are demanding documented adversarial testing. Most teams still don't have the tooling or methodology to meet that bar.
PROBE is a hands-on evaluation design jam where builders and researchers who care about AI safety + security come together to share and co-design evaluation methods. PROBE is designed as a discussion x working session designed to share and build benchmark datasets, stress-test harnesses, and open-source tooling that the community can reference as they stress test and evaluate agents.
The goal is to connect with other builders and researchers interested in AI safety and assurance - and create a working group for AI assurance.
Curated by AI Safety Node/Ecomonitor (www.aisafetynode.com) and hosted with Ojin(https://ojin.ai/).
**** Program and Working Group ****
Safeguard Stress Testing Design
Share and work on evaluation design that probe guardrails, test refusal mechanisms and results (if past evaluations are available)
Evaluation Infrastructure
Discuss evaluation tools and infrastructure. Share and work on reusable harnesses, metrics pipelines, and reporting tools that make adversarial testing repeatable, not one-off.
Schedule
18-19:00: Kick off and guest speaker on evaluation method
Speakers from AI Safety Hub in Berlin and independent research firms to discuss evaluation methods and red teaming x agentic harnesses.
19:00 - 20:00: Evaluation design sprint: evaluation design-> implementation sprint
20:00-20:30: Share eval design work through standup and researcher exchange
This event is open to all. Safety researchers building adversarial benchmarks and agent builders/ML engineers who are exploring eval tooling will find the evening especially relevant.