

"I Can't Assist With That": The AI Refusal Problem in Security
Security practitioners keep hitting the same wall: the commercial LLMs meant to speed up their work refuse the very prompts their job depends on. A vulnerability-ID task, a payload for an authorized engagement, a post-exploitation question — and back comes "I can't assist with that." Guardrails tuned to catch "write me malware" end up blocking real work.
In this webinar, Luca Mannini, AI Researcher at Cracken, digs into why this happens, how to measure it, and what to do about it.
What he will cover:
The Refusal Problem — Highlighting the practical friction users experience with commercial models refusing cyber-related prompts.
RedLineBench — Cracken's open benchmark concerning findings on how current models perform and where they trigger refusals during security and red-teaming tasks.
Domain-Specific Abliteration — Introducing domain-specific mitigation approaches as a way to bypass unnecessary restrictions and handle refusals effectively.
What you'll learn:
Why commercial models refuse legitimate cyber prompts, and where that friction hits hardest in real red-teaming work.
How RedLineBench scores refusal and capability separately, and what its findings reveal about how current models perform on security tasks.
Which models hold up on offensive-security work — who refuses, who complies, and who's actually useful when they do.
How domain-specific abliteration bypasses unnecessary restrictions, and when it's the right way to handle refusals.
Who should attend: red teamers, penetration testers, security engineers, and AI/ML practitioners working where offensive security meets LLMs.
Explore the benchmark: cracken-ai/redline-bench