arXiv:2602.04003cs.AI2026-02被引 3

攻击者通过伪造解释误导用户信任错误AI判断,威胁人机协作安全。

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making

  • 操纵大模型解释的表述方式,诱导用户误信错误结果。
  • 用户对恶意与正常解释的信任度几乎无差,错误信任率达90%以上。
  • 专家式表达最易欺骗,尤其在复杂任务和低教育背景人群中。

当前大多数对抗性威胁针对模型计算行为,而非依赖模型的人类。然而现代AI系统越来越多地嵌入人类决策流程,用户需解读并响应模型建议。大语言模型(LLMs)生成流畅自然语言解释,影响用户对AI输出的感知与信任,暴露出认知层面的新攻击面:AI与用户间的沟通通道。本文提出对抗性解释攻击(AEAs),即攻击者通过篡改LLM生成解释的表述框架,操控用户对错误输出的信任度。我们以‘信任偏差差距’为度量指标,揭示一种行为风险——即使模型预测错误,精心设计的解释仍可维持用户高信任。通过超过200名参与者的实验,系统考察解释框架的四个维度:推理模式、证据类型、沟通风格与呈现格式。结果显示,用户对恶意与正常解释的信任水平近乎一致,恶意解释在多数情况下保留了超过90%的正常信任。最脆弱情形出现在攻击解释模拟专家表达时,结合权威证据、中立语气与领域适配推理。该风险在高难度任务、事实驱动领域,以及教育程度较低、年龄较轻或高度信赖AI的群体中尤为突出。

原文摘要 · Abstract (English)

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and act on model recommendations. Large Language Models (LLMs) generate fluent natural-language explanations that shape how users perceive and trust AI outputs, revealing a new attack surface at the cognitive layer: the communication channel between AI and its users. We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Using this metric as a lens, we highlight a behavioral risk where persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong. To characterize this threat, we conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Our findings show that users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI.

对抗攻击人机信任大模型解释认知安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。