arXiv:2503.17645cs.AIcs.LG2025-03被引 1

用拼图数据集揭示大模型如何内部区分推理对错

A Modular Dataset to Demonstrate LLM Abstraction Capability

  • 设计结构化拼图数据集,自动验证每步推理正确性
  • 80%以上准确率证明模型能内部分辨正确与错误推理
  • 发现中间层编码抽象逻辑概念,适合研究模型可解释性

大型语言模型虽表现优异,但常因幻觉和逻辑谬误导致推理错误。为探究其内部推理表征,我们提出ArrangementPuzzle——一个具有结构化解法且支持自动化步骤级正确性验证的新型谜题数据集。在该数据集上训练分类器分析LLM激活值,结果表明其预测推理正确性的准确率超过80%,说明模型内部能区分正确与错误推理步骤,最强表征出现在Transformer中后期层。进一步分析显示,模型在中间激活层编码了抽象推理概念,能区分逻辑等价与语义等价。这些发现为理解大模型推理机制提供洞见,有助于提升AI可靠性与可解释性,实现对推理过程的调控与优化。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit impressive capabilities but struggle with reasoning errors due to hallucinations and flawed logic. To investigate their internal representations of reasoning, we introduce ArrangementPuzzle, a novel puzzle dataset with structured solutions and automated stepwise correctness verification. We trained a classifier model on LLM activations on this dataset and found that it achieved over 80% accuracy in predicting reasoning correctness, implying that LLMs internally distinguish between correct and incorrect reasoning steps, with the strongest representations in middle-late Transformer layers. Further analysis reveals that LLMs encode abstract reasoning concepts within the middle activation layers of the transformer architecture, distinguishing logical from semantic equivalence. These findings provide insights into LLM reasoning mechanisms and contribute to improving AI reliability and interpretability, thereby offering the possibility to manipulate and refine LLM reasoning.

大模型推理可解释性逻辑能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。