高权限大模型代理易被文档指令劫持,导致隐私数据泄露。
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
- 构建三维度伪造指令分类法,量化恶意指令隐蔽性。
- 实测代理数据外泄成功率高达85%,跨语言跨位置稳定存在。
- 用户与现有防御均无法有效识别,暴露严重安全缺口。
具备高权限的大型语言模型代理正被广泛用于自动化处理外部文档,执行项目指令,但其拥有终端访问、文件系统控制及网络连接权,却缺乏有效安全监管。我们识别并系统测量了这一信任模型的根本性漏洞,称为‘可信执行困境’:代理会以高频率执行嵌入文档中的指令,包括恶意指令,因其无法区分恶意指令与合法配置信息。此漏洞是指令遵循设计范式的结构性产物,而非实现缺陷。为此,我们提出三维分类体系(语言伪装、结构混淆、语义抽象),并构建了包含500个真实世界README文件的基准测试集ReadSecBench,支持可复现评估。在商用计算机使用代理上实验显示,端到端数据外泄成功率最高达85%,且在五种编程语言和三种注入位置下均保持一致。跨模型评估在四个主流大模型家族中验证,语义合规性表现一致。15名参与者的用户研究显示检测率为0%;对12种规则型与6种基于大模型的防御机制评估表明,两类方法均无法在不产生高误报率的情况下实现可靠检测。综合结果揭示出代理在功能合规与安全意识之间存在持续性的‘语义安全鸿沟’,证明文档嵌入式指令注入是当前高权限大模型代理部署中持久且未被缓解的重大威胁。
原文摘要 · Abstract (English)
High-privilege LLM agents that autonomously process external documentation are increasingly trusted to automate tasks by reading and executing project instructions, yet they are granted terminal access, filesystem control, and outbound network connectivity with minimal security oversight. We identify and systematically measure a fundamental vulnerability in this trust model, which we term the \emph{Trusted Executor Dilemma}: agents execute documentation-embedded instructions, including adversarial ones, at high rates because they cannot distinguish malicious directives from legitimate setup guidance. This vulnerability is a structural consequence of the instruction-following design paradigm, not an implementation bug. To structure our measurement, we formalize a three-dimensional taxonomy covering linguistic disguise, structural obfuscation, and semantic abstraction, and construct \textbf{ReadSecBench}, a benchmark of 500 real-world README files enabling reproducible evaluation. Experiments on the commercially deployed computer-use agent show end-to-end exfiltration success rates up to 85\%, consistent across five programming languages and three injection positions. Cross-model evaluation on four LLM families in a simulation environment confirms that semantic compliance with injected instructions is consistent across model families. A 15-participant user study yields a 0\% detection rate across all participants, and evaluation of 12 rule-based and 6 LLM-based defenses shows neither category achieves reliable detection without unacceptable false-positive rates. Together, these results quantify a persistent \emph{Semantic-Safety Gap} between agents' functional compliance and their security awareness, establishing that documentation-embedded instruction injection is a persistent and currently unmitigated threat to high-privilege LLM agent deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。