构建安全沙箱系统,评估人机交互中AI代理的风险表现
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
- 设计模块化沙箱环境,模拟多轮人机交互与工具使用
- 1840次实验显示主流大模型在超50%场景存在安全风险
- 适合关注AI安全、伦理及可控性的研究人员与开发者
随着AI代理在与人类用户和工具的交互中自主性增强,交互安全风险日益突出。我们提出HAICOSYSTEM,一个用于评估复杂社会互动中AI代理安全性的框架。该系统包含模块化沙箱环境,可模拟人类用户与AI代理之间的多轮交互,其中AI代理配备多种工具(如患者管理平台),应对多样化场景(如用户试图访问他人病历)。为全面评估安全性,我们构建涵盖操作、内容、社会与法律风险的多维评价体系。基于92个场景、7个领域(如医疗、金融、教育)的1840次模拟实验表明,该系统能有效复现真实的人机交互与工具使用。实验发现,当前主流的大语言模型(包括商用与开源版本)在超过50%的情况下暴露安全风险,且面对模拟恶意用户时风险更高。研究揭示了构建可安全应对复杂交互的AI代理仍面临重大挑战。为推动AI代理安全生态建设,我们开源代码平台,支持用户自定义场景、模拟交互并评估代理的安全性与性能。
原文摘要 · Abstract (English)
AI agents are increasingly autonomous in their interactions with human users and tools, leading to increased interactional safety risks. We present HAICOSYSTEM, a framework examining AI agent safety within diverse and complex social interactions. HAICOSYSTEM features a modular sandbox environment that simulates multi-turn interactions between human users and AI agents, where the AI agents are equipped with a variety of tools (e.g., patient management platforms) to navigate diverse scenarios (e.g., a user attempting to access other patients' profiles). To examine the safety of AI agents in these interactions, we develop a comprehensive multi-dimensional evaluation framework that uses metrics covering operational, content-related, societal, and legal risks. Through running 1840 simulations based on 92 scenarios across seven domains (e.g., healthcare, finance, education), we demonstrate that HAICOSYSTEM can emulate realistic user-AI interactions and complex tool use by AI agents. Our experiments show that state-of-the-art LLMs, both proprietary and open-sourced, exhibit safety risks in over 50\% cases, with models generally showing higher risks when interacting with simulated malicious users. Our findings highlight the ongoing challenge of building agents that can safely navigate complex interactions, particularly when faced with malicious users. To foster the AI agent safety ecosystem, we release a code platform that allows practitioners to create custom scenarios, simulate interactions, and evaluate the safety and performance of their agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。