

Evaluating AI That Outgrows Our Tests: Frontier Model Risks and Behaviour
Understanding how a frontier AI model behaves—and what kinds of dangerous behavior might emerge—is one of the most difficult challenges in the field today. No single benchmark captures the full range of ways a model could go wrong. This raises a fundamental question: How do we characterise specific models, one at a time, with the methods we have, and where do current methods fall short?
Clement Neo, Founder and Research Lead at Neo Research, joins us to tackle this question head-on. He will draw from the lab's groundbreaking new report, recently featured in the South China Morning Post, which reveals that Chinese frontier models are rapidly developing "evaluation awareness"—the ability to recognise when they are being tested and adjust their behavior accordingly.
This phenomenon has profound implications for AI safety. As Clement Neo explains, "It would mean that whatever testing the model developers themselves do might not reflect the actual behaviour of a model once it gets deployed. And that’s a really big problem".
In this talk, Clement will cover:
The State of Evaluation: How independent evaluation approaches the challenge through testing adversarial robustness, honesty under pressure, evaluation awareness, and behavior in realistic agentic environments.
The Growing Gap: Why this kind of evaluation is becoming harder as models grow more agentic and existing benchmarks saturate, leading to issues like "alignment faking" and "sandbagging" .
The Neo Research Findings: A deep dive into specific results, including how Moonshot AI's Kimi K2.6 exhibited evaluation awareness in 60% of test instances, and how DeepSeek's V4 Pro, while scoring lower (17%), demonstrated in its chain-of-thought that it recognised it was in a fictional test scenario.
The Path Forward: What it will take to keep characterising model behaviour reliably, especially as the capabilities gap between Chinese and Western models continues to close.
This is a critical conversation for anyone concerned with the future of AI governance, safety, and the reliability of the tests we rely on to certify frontier systems.
Who should attend: This talk is essential for AI safety researchers, developers, and policy professionals, but it is also designed to be accessible to anyone with a stake in the future of AI, including non-technical audiences. While we will discuss technical evaluations, the core questions—how do we know if a model is safe, can we trust our tests, and what happens when models learn to hide their true behaviour—are fundamentally about trust, governance, and decision-making. No prior technical background is required to follow the key insights or to understand why this matters for all of us.
About the Speaker:
Clement is the Founder and Research Lead of Neo Research (新衡), a Singapore-based independent AI safety research and evaluation organisation. Neo Research evaluates frontier AI models across a range of safety-relevant behaviours, including deception and agentic risk. Their recent evaluation of DeepSeek v4 Pro received coverage from the South China Morning Post.
Website: neoresearch.ai
Neo Research is hiring! neoresearch.ai/careers
Stay connected with AI Safety Hong Kong.
LinkedIn: https://www.linkedin.com/company/ai-safety-hong-kong/
Substack: http://substack.com/@aisafetyhk
Website: https://www.aisafetyhk.org/