arXiv:2601.08333cs.AI2026-01被引 3

AI代理架构中的信息信任机制存在根本缺陷,导致错误结论被误认为合理。

Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant

  • 用'语义清洗'描述代理架构中信任与证据混淆的系统性问题
  • 证明在标准设计下,循环论证不可避免,无法通过模型改进消除
  • 适合研究AI可信推理、认知架构或形式化验证的研究者

基于大语言模型的智能体架构会系统性地混淆信息传递机制与知识确证机制。我们将其定义为语义清洗:当命题缺乏或证据薄弱时,仍可通过架构中受信任的接口被系统接纳为有效。该现象构成格蒂尔问题的架构实现——命题获得高知识地位,但其理由与真实性之间无实际关联。与经典格蒂尔案例不同,这种效应并非偶然,而是由架构决定且可系统复现。核心结论是‘必然自许可定理’:在标准架构假设下,循环确证无法消除。我们提出确证侵蚀原理作为根本解释,并表明缩放、模型优化及大模型评判方案在类型层面均无法根除此问题。

原文摘要 · Abstract (English)

LLM-based agent architectures systematically conflate information transport mechanisms with epistemic justification mechanisms. We formalize this class of architectural failures as semantic laundering: a pattern where propositions with absent or weak warrant are accepted by the system as admissible by crossing architecturally trusted interfaces. We show that semantic laundering constitutes an architectural realization of the Gettier problem: propositions acquire high epistemic status without a connection between their justification and what makes them true. Unlike classical Gettier cases, this effect is not accidental; it is architecturally determined and systematically reproducible. The central result is the Theorem of Inevitable Self-Licensing: under standard architectural assumptions, circular epistemic justification cannot be eliminated. We introduce the Warrant Erosion Principle as the fundamental explanation for this effect and show that scaling, model improvement, and LLM-as-judge schemes are structurally incapable of eliminating a problem that exists at the type level.

AI可信性认知架构语义清洗确证理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。