测试智能体对话中的隐私与安全风险,发现主流模型漏洞率超60%。
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
- 构建多轮动态对话基准,模拟真实场景下恶意请求嵌入
- 88%隐私攻击成功,60%安全攻击导致工具滥用或偏好篡改
- 首次将隐私与安全统一在多智能体交互中评估,适合安全研究者
随着语言模型演变为能代表用户行动和沟通的自主智能体,多智能体生态系统的安全性成为核心挑战。个人助手与外部服务提供者之间的互动,在效率与防护之间存在根本矛盾:有效协作需信息共享,但每次交流都可能暴露新攻击面。我们提出ConVerse,一个用于评估智能体间交互中隐私与安全风险的动态基准。该基准覆盖旅行、房地产、保险三个实际领域,包含12种用户角色及超过864个上下文相关的攻击(其中611项为隐私攻击,253项为安全攻击)。与以往单智能体设置不同,ConVerse模拟自主、多轮的智能体对话语境,将恶意请求嵌入合理对话中。隐私评估采用三级分类体系,衡量抽象质量;安全攻击则针对工具使用和偏好操纵。对七种前沿模型的评估显示,隐私攻击成功率高达88%,安全漏洞率达60%,且越强模型泄漏越多。通过在交互式多智能体环境中统一隐私与安全,ConVerse将安全重新定义为通信的涌现属性。
原文摘要 · Abstract (English)
As language models evolve into autonomous agents that act and communicate on behalf of users, ensuring safety in multi-agent ecosystems becomes a central challenge. Interactions between personal assistants and external service providers expose a core tension between utility and protection: effective collaboration requires information sharing, yet every exchange creates new attack surfaces. We introduce ConVerse, a dynamic benchmark for evaluating privacy and security risks in agent-agent interactions. ConVerse spans three practical domains (travel, real estate, insurance) with 12 user personas and over 864 contextually grounded attacks (611 privacy, 253 security). Unlike prior single-agent settings, it models autonomous, multi-turn agent-to-agent conversations where malicious requests are embedded within plausible discourse. Privacy is tested through a three-tier taxonomy assessing abstraction quality, while security attacks target tool use and preference manipulation. Evaluating seven state-of-the-art models reveals persistent vulnerabilities; privacy attacks succeed in up to 88% of cases and security breaches in up to 60%, with stronger models leaking more. By unifying privacy and security within interactive multi-agent contexts, ConVerse reframes safety as an emergent property of communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。