攻击者通过改变大模型推理风格,让其产生错误判断却不篡改内容。
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
- 用生成式风格注入技术,将检索内容改写为拖延或急躁的病态语气。
- 在HotpotQA和FEVER上使推理步骤最多增加4.4倍,导致提前出错。
- 提出实时监控系统,通过分析推理过程发现异常风格变化。
依赖外部检索的大语言模型代理正被广泛部署于高风险场景。现有攻击多聚焦于内容伪造或指令注入,我们识别出一种新型过程性攻击面:代理的推理风格。提出推理风格污染(RSP)范式,通过操纵信息处理方式而非内容本身实现攻击。引入生成式风格注入(GSI)方法,将检索文档重写为“分析瘫痪”或“认知仓促”等病态语调,不改变事实且无需显式触发词。为量化此类转变,构建推理风格向量(RSV),追踪验证深度、自信心与注意力焦点。在HotpotQA与FEVER数据集上,基于ReAct、Reflection及树状思维(ToT)架构的实验表明,GSI显著降低性能:推理步数最多增加4.4倍,或引发过早错误,并成功绕过先进内容过滤机制。最后提出RSP-M,一种轻量级运行时监控器,可实时计算RSV指标并在超过安全阈值时触发警报。本工作证明推理风格是独立且可利用的漏洞,需建立超越静态内容分析的过程级防御体系。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents relying on external retrieval are increasingly deployed in high-stakes environments. While existing adversarial attacks primarily focus on content falsification or instruction injection, we identify a novel, process-oriented attack surface: the agent's reasoning style. We propose Reasoning-Style Poisoning (RSP), a paradigm that manipulates how agents process information rather than what they process. We introduce Generative Style Injection (GSI), an attack method that rewrites retrieved documents into pathological tones--specifically "analysis paralysis" or "cognitive haste"--without altering underlying facts or using explicit triggers. To quantify these shifts, we develop the Reasoning Style Vector (RSV), a metric tracking Verification depth, Self-confidence, and Attention focus. Experiments on HotpotQA and FEVER using ReAct, Reflection, and Tree of Thoughts (ToT) architectures reveal that GSI significantly degrades performance. It increases reasoning steps by up to 4.4 times or induces premature errors, successfully bypassing state-of-the-art content filters. Finally, we propose RSP-M, a lightweight runtime monitor that calculates RSV metrics in real-time and triggers alerts when values exceed safety thresholds. Our work demonstrates that reasoning style is a distinct, exploitable vulnerability, necessitating process-level defenses beyond static content analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。