arXiv:2601.05478cs.CL2026-01被引 2

发现大模型易被精心伪造的证据误导,提出防御机制。

The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence

  • 用多角色大模型协作生成难以辨伪的虚假证据
  • 模型对这类证据信念值平均提升93%,推荐严重失真
  • 提出欺骗意图预警机制,可有效降低误信风险

为可靠辅助人类决策,大模型需在面对误导信息时保持事实信念。尽管现有模型能抵抗直接错误信息,我们发现其对复杂、难辨伪的证据存在根本性脆弱性。为此,我们提出MisBelief框架,通过多角色大模型多轮协作生成逻辑自洽但事实错误的论据,模拟渐进式说服过程。该方法生成4800个实例,覆盖三个难度等级,评估7种代表性大模型。结果表明:模型对直接虚假信息有抵抗力,但对经精心打磨的证据极为敏感——虚假信念得分平均上升93.0%,严重损害下游决策。为此,我们提出欺骗意图屏蔽(Deceptive Intent Shielding, DIS)机制,通过推断证据背后的欺骗意图提供早期预警。实证显示,DIS能持续缓解信念偏移,促使模型更审慎评估证据。

原文摘要 · Abstract (English)

To reliably assist human decision-making, LLMs must maintain factual internal beliefs against misleading injections. While current models resist explicit misinformation, we uncover a fundamental vulnerability to sophisticated, hard-to-falsify evidence. To systematically probe this weakness, we introduce MisBelief, a framework that generates misleading evidence via collaborative, multi-round interactions among multi-role LLMs. This process mimics subtle, defeasible reasoning and progressive refinement to create logically persuasive yet factually deceptive claims. Using MisBelief, we generate 4,800 instances across three difficulty levels to evaluate 7 representative LLMs. Results indicate that while models are robust to direct misinformation, they are highly sensitive to this refined evidence: belief scores in falsehoods increase by an average of 93.0\%, fundamentally compromising downstream recommendations. To address this, we propose Deceptive Intent Shielding (DIS), a governance mechanism that provides an early warning signal by inferring the deceptive intent behind evidence. Empirical results demonstrate that DIS consistently mitigates belief shifts and promotes more cautious evidence evaluation.

大模型安全欺骗检测信念修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。