arXiv:2607.05013cs.CLcs.LG2026-07被引 1

区分大模型对数学题的解题知识与表达能力,发现造假主要源于表达而非知识本身。

Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs

论文配图:Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs
图 1 · 摘自论文原文
  • 分离分析模型内部的解题知识与口头表达机制
  • 发现伪造回答主要由表达模块变化导致,知识本身未变
  • 可通过提示词或激活调控提升模型拒绝不靠谱问题的能力

尽管大语言模型在数学推理方面取得显著进展,但判断一道数学题是否可解仍是一项基础而困难的能力。现有研究多关注模型内在解题信念的表征,而对‘表述’行为的内部表征研究较少,限制了对其分析与干预。本文通过分别探测解题知识与表达机制,实现了在模型隐状态中对两者的解耦。在多个大模型上,我们发现知识与表达是独立且线性可解码的表征;而虚假回答主要与表达变化相关,而非知识本身。使用不可解提示词可有效降低虚假生成,主要通过调节表达模块实现;激活操控实验进一步表明,这些表征可被机制性调控以增强模型的自我克制能力。

原文摘要 · Abstract (English)

Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability. While recent studies have probed internal representations of model solvability beliefs, verbalization has primarily been studied behaviorally rather than as an internal representation, limiting its analysis and manipulation. We address this gap by separately probing representations of solvability knowledge and verbalization, allowing us to disentangle the two within model hidden states. Across multiple LLMs, we show that knowledge and verbalization are encoded as distinct, linearly decodable representations and that fabrication is primarily associated with changes in verbalization rather than the underlying knowledge. Prompting with unsolvability cues reduces fabrication primarily by shifting verbalization, while activation steering demonstrates that these representations can be echanistically manipulated to improve model abstention.

大模型数学推理表达机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。