Operational Safety for Language Model Agents
Abstract
As language model agents gain access to tools, external services, persistent context, and increasingly autonomous workflows, safety depends not only on model behavior but also on the systems in which agents operate. This social will bring together researchers and practitioners working on agent safety, language model evaluation, security, monitoring, and deployment to discuss operational approaches for safely running increasingly capable agents. Topics may include sandboxing, permissions and access control, monitoring of agent actions, human approval mechanisms, long-horizon execution, recovery from unexpected behavior, incident detection, and post-deployment evaluation. Recent cases - including OpenAI models escaping intended isolation and reaching Hugging Face systems, Claude models accessing real third-party infrastructure during evaluations, and Hacktron using Claude to help develop an exploit chain that reached OpenAI internal repositories - highlight the growing importance of operational controls around capable agents. The event will combine short invited talks with moderated and audience discussion, with the goal of connecting research on model behavior with practical questions around deployment and oversight.
Log in and register to view live content
| COLM uses cookies for essential functions only. We do not sell your personal information. Our Privacy Policy » |