

How Do You Build a Benchmark That Agents Can't Outsmart?
DeepSeek V4.1 Flash scored 11 out of 11 on our AI hacking benchmark for $4.65. Then we audited every run. Six followed the attack path we designed. Five used routes we never intended to test, and our scoring rules could not tell the difference.
The model did not cheat. It did what we asked and found the fastest way to do it. A hacking agent can find flaws in more than code. It can find them in the environment, the grader, or the definition of success. Ours broke on the third.
For security engineers, AI and ML teams, and anyone who builds, buys, or cites model evaluations.
We'll cover:
Public cases where agents beat the eval instead of the task
What we changed in the benchmark, and what we wish we had considered before we built it
What to ask the next time someone shows you a benchmark number
30 minute session with Yanir Tsarimi, Co-founder & CPO at Enclave.