arXiv:2510.14925cs.AIcs.CL2025-10

大模型自信错误可能并非偶然,而是稳定存在的认知陷阱。

False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs

  • 用康德式承诺门和线性反馈模型,揭示错误可自洽且稳定。
  • 实测显示错误项与正确项在局部敏感度上无显著差异。
  • 适合关注模型可信度与内在逻辑缺陷的研究者阅读。

大语言模型中的高置信度错误常被视为脆弱的失败。我们提出另一种可能:某些错误可能是虚假固定点,即局部稳定、内部自洽且自信错误的状态。这使鲁棒性与真实性追踪相分离。通过康德式承诺门框架和最小线性反馈模型,我们展示了稳定性与正确性可分离。在三个开源模型中,通过隐藏状态敏感性探针发现,过度自信的错误项在局部脆弱性上并不显著高于自信正确的项。采用回避意识的自我批判可减少过度自信的错误承诺,但以降低覆盖率为代价;而基于规则的显式反馈门C3-R虽能优化权衡,却无法彻底消除该矛盾。这些结果支持但未证实,高信噪比惯性与表征压缩是导致稳定误校准的潜在机制。

原文摘要 · Abstract (English)

High-confidence errors in large language models are often treated as fragile failures. We study an alternative: some errors may be false fixed points, locally stable, internally coherent, and confidently wrong. This separates robustness from truth-tracking. We develop the separation through a Kantian commitment-gate framing and a minimal linear feedback model in which stability and correctness can diverge. Across three open-weight models, overconfident wrong items are not systematically more locally fragile than confidently correct items under our hidden-state sensitivity probes. Abstention-aware self-critique reduces overconfident wrong commitments by sacrificing coverage, and C3-R, a rule-based explicit feedback gate, sharpens that tradeoff rather than eliminating it. These results motivate, but do not establish, high signal-to-noise (high-SNR) inertia and representational compression as possible mechanisms for stable miscalibration.

大模型误差认知稳定可信度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。