世界模型可能被利用,导致智能体执行恶意命令或泄露数据。
False Prophets: On the Security of World Models in Agentic Systems
- 发现世界模型特有漏洞,可被攻击者利用
- 攻击成功率高达95%,可引发命令执行等危害
- 适合关注智能体安全的开发者和研究者
大型语言模型如今驱动自主智能体在不同环境中执行复杂多步任务。准确可靠的执行依赖于智能体对自身行为结果的预测能力。近期研究提出通过专门训练的环境模拟器——世界模型来增强预测能力。尽管世界模型能提升性能,也可能误导智能体执行有害操作,带来严重安全与隐私风险。本文揭示了智能体系统中世界模型存在的多种安全漏洞,证明攻击者可在基于终端的智能体中执行恶意代码或提取敏感信息。为推动后续研究,我们构建了一个面向文本型世界模型的安全基准数据集。我们认为部分风险源于近似世界建模的本质,展示了攻击者可使智能体流水线产生误判,成功率高达95%,可能导致意外命令执行、服务拒绝、钱包耗尽及私密信息窃取。最后,我们为从业者提供实用建议,以缓解已发现的危害并强化智能体系统安全性。
原文摘要 · Abstract (English)
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained environment simulators-world models. While world models can improve performance, they can also mislead agents into executing harmful actions, creating significant security and privacy risks. In this paper, we raise security concerns regarding the usage of world models in agentic systems. We discover a range of world model specific vulnerabilities, which can be exploited in terminal-based agents to execute malicious code or extract sensitive data. To facilitate future development, we introduce a security benchmark dataset designed for text-based world models. We argue that some risks are intrinsic to approximate world modeling, and show that attackers can induce mispredictions in agentic pipelines with up to 95% success rate, possibly resulting in unintended command execution, denial of service, drainage of wallet and private information extraction. Finally, we provide practical recommendations for practitioners to mitigate the discovered harms and harden agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。