攻击深度研究智能体的思维链,用伪造文档污染报告生成
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
- 通过伪造文档与链式协调实现任务规划劫持
- 5个注入文档使报告污染率达26.4%(PRISM指标)
- 适合研究安全、大模型对抗防御的学者关注
深度研究智能体将开放问题分解为子任务,多轮检索网络证据并合成长篇报告。这一流程在规划层形成污染面:恶意文档进入检索池后可引导后续问题,将局部注入转化为报告级污染。本文提出FORGE(用于智能体利用的伪造协同推理链),一种两级攻击方法,结合文档内推理伪造与文档间链路协调,劫持子任务规划。我们进一步引入PRISM度量,按认知类型加权受感染报告内容;提出根查询锚定(RQA)轻量级防御,将递归生成绑定至初始查询。在25个查询上,网络版FORGE在注入5个文档时达到26.4%的PRISM值,并出现深度迁移现象——中毒内容从明显表述转移到事实前提中。在10个防御测试查询上,RQA将PRISM从38.5%降至18.3%。
原文摘要 · Abstract (English)
Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-layer poisoning surface: adversarial documents that enter the retrieval pool can steer follow-up questions and turn a local injection into report-level contamination. We present FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation), a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning. We further introduce the PRISM metric, which weights infected report claims by cognitive type, and Root Query Anchoring, a lightweight defense that ties recursive follow-up generation to the root query. Across 25 queries, Network FORGE reaches 26.4% PRISM with five injected documents and exhibits depth migration, in which recursive synthesis shifts poisoned content from overt framing into factual premises. On the 10-query defense subset, RQA (Root Query Anchoring) reduces PRISM from 38.5% to 18.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。