arXiv:2507.06134cs.AI2025-07中稿 · ICLR被引 63

评估真实世界AI代理安全性的综合框架,涵盖8类风险。

OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety

  • 用真实工具和多轮任务评估代理行为,突破模拟环境限制。
  • 5个主流大模型在安全敏感任务中不安全行为占比达51.2%~72.7%。
  • 支持快速扩展工具与攻击策略,适合安全研究者使用。

具备解决复杂日常任务能力的AI代理已可部署于真实场景,但其潜在不安全行为亟需严格评估。现有基准大多依赖模拟环境、局限任务域或不真实的工具抽象。我们提出OpenAgentSafety,一个全面且模块化的评估框架,覆盖八类关键风险。不同于以往工作,该框架评估与真实工具(如浏览器、代码执行环境、文件系统、bash shell、消息平台)交互的代理,并支持350多个多轮、多用户任务,涵盖良性与对抗性用户意图。框架具备可扩展性,研究者可轻松添加工具、任务、网站及对抗策略。结合规则分析与大模型评判,能检测显性和隐性不安全行为。对五个主流大模型在代理场景中的实证分析显示,在安全敏感任务中不安全行为占比从Claude-Sonnet-3.7的51.2%到o3-mini的72.7%,凸显严重安全隐患,亟需更强防护措施方可部署。

原文摘要 · Abstract (English)

Recent advances in AI agents capable of solving complex, everyday tasks, from scheduling to customer service, have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous evaluation. While prior benchmarks have attempted to assess agent safety, most fall short by relying on simulated environments, narrow task domains, or unrealistic tool abstractions. We introduce OpenAgentSafety, a comprehensive and modular framework for evaluating agent behavior across eight critical risk categories. Unlike prior work, our framework evaluates agents that interact with real tools, including web browsers, code execution environments, file systems, bash shells, and messaging platforms; and supports over 350 multi-turn, multi-user tasks spanning both benign and adversarial user intents. OpenAgentSafety is designed for extensibility, allowing researchers to add tools, tasks, websites, and adversarial strategies with minimal effort. It combines rule-based analysis with LLM-as-judge assessments to detect both overt and subtle unsafe behaviors. Empirical analysis of five prominent LLMs in agentic scenarios reveals unsafe behavior in 51.2% of safety-vulnerable tasks with Claude-Sonnet-3.7, to 72.7% with o3-mini, highlighting critical safety vulnerabilities and the need for stronger safeguards before real-world deployment.

AI安全评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。