Cover Image for How Do You Build a Benchmark That Agents Can't Outsmart?
Cover Image for How Do You Build a Benchmark That Agents Can't Outsmart?
Avatar for Enclave
Presented by
Enclave
The autonomous security team for the enterprise. Enclave’s agents find vulnerabilities, prepare fixes, and retest the attack path after deployment.
Hosted By
Private Event

How Do You Build a Benchmark That Agents Can't Outsmart?

Virtual
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

​DeepSeek V4.1 Flash scored 11 out of 11 on our AI hacking benchmark for $4.65. Then we audited every run. Six followed the attack path we designed. Five used routes we never intended to test, and our scoring rules could not tell the difference.

​The model did not cheat. It did what we asked and found the fastest way to do it. A hacking agent can find flaws in more than code. It can find them in the environment, the grader, or the definition of success. Ours broke on the third.

​For security engineers, AI and ML teams, and anyone who builds, buys, or cites model evaluations.

​We'll cover:

  • ​Public cases where agents beat the eval instead of the task

  • ​What we changed in the benchmark, and what we wish we had considered before we built it

  • ​What to ask the next time someone shows you a benchmark number

​30 minute session with Yanir Tsarimi, Co-founder & CPO at Enclave.

Avatar for Enclave
Presented by
Enclave
The autonomous security team for the enterprise. Enclave’s agents find vulnerabilities, prepare fixes, and retest the attack path after deployment.
Hosted By