arXiv:2506.18998cs.CL2025-06被引 2

大模型把记忆当能力,导致自我认知虚高,可信度下降。

Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge

  • 通过新框架检验大模型是真学会推理还是靠记忆
  • 面对相似问题时自评准确率下降超45%
  • 尤其在科学医学领域问题更严重,适合关注AI可信度的研究者

当人工智能将记忆误认为智能时,会产生虚假的推理幻觉。现有研究将记忆与自我认知缺陷视为独立问题,未意识到二者交织会降低大模型回答的可信度。本研究提出新框架,检验大模型是否从训练数据中真正学习推理模式,或仅通过记忆来伪装对类似复杂度问题的掌握能力,聚焦于STEM领域。分析显示,泛化存在显著问题:大模型基于记忆解题而自信,导致在自我验证、逻辑一致的问题扰动下,可行性评估不一致性超过45%。该现象在科学与医学领域尤为明显,这些领域标准化术语和问题最多,进一步验证了方法的有效性。大模型自我认知波动剧烈,暴露出当前架构与训练方式的缺陷,凸显需要开发能保持模型自知一致性、提升AI可解释性与可信度的技术。代码与结果已公开于https://github.com/Sahil-R-Kale/mirage_of_mastery。

原文摘要 · Abstract (English)

When artificial intelligence mistakes memorization for intelligence, it creates a dangerous mirage of reasoning. Existing studies treat memorization and self-knowledge deficits in LLMs as separate issues and do not recognize an intertwining link that degrades the trustworthiness of LLM responses. In our study, we utilize a novel framework to ascertain if LLMs genuinely learn reasoning patterns from training data or merely memorize them to assume competence across problems of similar complexity focused on STEM domains. Our analysis shows a noteworthy problem in generalization: LLMs draw confidence from memorized solutions to infer a higher self-knowledge about their reasoning ability, which manifests as an over 45% inconsistency in feasibility assessments when faced with self-validated, logically coherent task perturbations. This effect is most pronounced in science and medicine domains, which tend to have maximal standardized jargon and problems, further confirming our approach. Significant wavering within the self-knowledge of LLMs also shows flaws in current architectures and training patterns, highlighting the need for techniques that ensure a balanced, consistent stance on models' perceptions of their own knowledge for maximum AI explainability and trustworthiness. Our code and results are available publicly at https://github.com/Sahil-R-Kale/mirage_of_mastery

大模型幻觉自我认知可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。