arXiv:2601.07085cs.HCcs.AI2026-01被引 7

大模型可能通过伪装成可信信息源,绕过人类的批判性思维机制。

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

  • 利用流畅性、帮助性等看似真实但低成本的特征,诱导人类信任
  • 认知负荷转移使用户放弃自主评估,更易受误导
  • 高智商用户反而更易被影响,挑战直觉认知

基于大语言模型(LLM)的对话式AI系统对人类认知构成新挑战,现有虚假信息与说服理论难以应对。本文提出‘认知特洛伊木马’假说:这些系统在优化过程中生成的自然流畅、乐于助人等特征,虽为真实表现,却缺乏人类语境下的信号价值——因人类产生此类特征成本高昂,而大模型可轻松实现。四种潜在绕过机制包括:流畅性与理解脱钩、信任与能力呈现无实际代价、将评估任务交由AI自身、以及优化过程催生讨好行为。该框架提出可验证预测,例如高认知能力用户可能更易受影响。这将AI安全重新定义为评估校准问题,而非仅防欺骗。

原文摘要 · Abstract (English)

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant epistemic risk from conversational AI may lie not in inaccuracy or intentional deception, but in something more fundamental: these systems may be configured, through optimization processes that make them useful, to present characteristics that bypass the cognitive mechanisms humans evolved to evaluate incoming information. The Cognitive Trojan Horse hypothesis draws on Sperber and colleagues' theory of epistemic vigilance -- the parallel cognitive process monitoring communicated information for reasons to doubt -- and proposes that LLM-based systems present 'honest non-signals': genuine characteristics (fluency, helpfulness, apparent disinterest) that fail to carry the information equivalent human characteristics would carry, because in humans these are costly to produce while in LLMs they are computationally trivial. Four mechanisms of potential bypass are identified: processing fluency decoupled from understanding, trust-competence presentation without corresponding stakes, cognitive offloading that delegates evaluation itself to the AI, and optimization dynamics that systematically produce sycophancy. The framework generates testable predictions, including a counterintuitive speculation that cognitively sophisticated users may be more vulnerable to AI-mediated epistemic influence. This reframes AI safety as partly a problem of calibration -- aligning human evaluative responses with the actual epistemic status of AI-generated content -- rather than solely a problem of preventing deception.

认知安全大模型风险信息可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。