

COLM Social: Operational Safety for Language Model Agents
COLM Social: Operational Safety
As AI agents move from answering questions to taking actions, safety becomes an operational problem. Agents can access tools, modify code, handle private data, and make decisions on a user’s behalf — raising practical questions around prompt injection, monitoring, permissions, authorization, and control.
This social brings together researchers working on how to make deployed agents safer: how to prevent agents from being hijacked, detect unsafe behavior as it happens, and make sure agents act only when and how users intended.
Agenda
Welcome
Talk · Sizhe Chen
Talk · Han Wang
Talk · Lily Chen
Talk · Thibaud Gloaguen
Panel discussion and audience Q&A
Wrap-up
Talks
Sizhe Chen · PhD student at UC Berkeley
Works on agent security and prompt-injection defenses, including StruQ and SecAlign.Han Wang · PhD candidate at UIUC
Han’s work includes MonitorBench, a benchmark for evaluating monitoring models through agents’ reasoning traces.Yen-Shan (Lily) Chen · Data Scientist at CyCraft Technology and researcher at National Taiwan University
Works on LLM and agent security, including TraceSafe, a benchmark for evaluating how well guardrails detect risks across multi-step tool-calling trajectories.Thibaud Gloaguen · PhD student at ETH Zurich
Thibaud’s work examines an important problem in coding agents: agents taking actions or making changes that users did not ask them to perform.
Panel
Eugene Bagdasarian · Assistant Professor at UMass Amherst and part-time at Google
Works on contextual agent security, policy enforcement, and systems including AirGapAgent, bringing perspectives from both research and production deployments.Seraphina Goldfarb-Tarrant · Head of AI Safety at Cohere
Leads AI safety work at Cohere, with experience spanning model safety, security, user simulation, and the practical decisions involved in deploying models to enterprise users.Zhiping Zhang · PhD student at Northeastern University
First author of AuthorizationBench, studying whether agents remain within the scope of actions users actually authorized.
Themes
The discussion will focus on two closely related questions:
How do we prevent unsafe actions?
Prompt injection, permissions, authorization, and mechanisms for constraining what agents are allowed to do.
How do we know when something is going wrong?
Monitoring agent behavior and reasoning, detecting failures during deployment, and deciding when human intervention is needed.
Organizers
Tatia Tsmindashvili
Raphael Kalandadze
Good to know
Talks will be approximately 10 minutes each.
Talks are intended to be accessible to a mixed research audience.
The talks will be followed by a broader panel discussion and audience Q&A.