arXiv:2601.06160cs.AI2026-01ACL被引 2

通过几何探测提升大模型数学推理的多样性,解决错误思路重复问题。

Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration

  • 用弱模型作为正交探针,向主空间补集注入异质推理信号。
  • 在数学基准上准确率提升62.4%,采样效率提高113.7%。
  • 适合需要增强推理多样性的大模型研究与应用者。

大语言模型在复杂数学推理任务中常出现‘推理坍缩’,即随机采样产生相同错误逻辑的词汇变体,而非真正语义上的探索。我们观察到失败的推理路径往往对应于模型隐藏状态几何中的低秩偏差流形,导致难以朝修正方向探索。为此,提出谱正交探索(SOE)框架,基于‘学生引导教师’范式:不采用弱辅助代理进行模仿,而是将其作为正交探针,向教师主空间的正交补集引入语义异构的推理信号。该干预使教师走向更丰富的推理轨迹,突破标准采样的局限。在数学基准上的实验表明,相较于基线方法,SOE平均准确率提升62.4%,平均采样效率提高113.7%,表明几何干预可有效缓解数学推理中的推理坍缩。我们还提供了初步证据,显示SOE在逻辑推理与代码生成任务中亦具有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often suffer from ''Reasoning Collapse'' on challenging mathematical reasoning tasks, where stochastic sampling produces lexical variations of the same erroneous logic rather than genuine semantic exploration. We observe that failed reasoning traces are often associated with a low-rank bias manifold in the model's hidden-state geometry, which reduces exploration toward corrective solution directions. To address this, we propose Spectral Orthogonal Exploration (SOE), a geometric inference framework under a ''Student Guides Teacher'' paradigm. Instead of using a weak auxiliary agent for imitation, SOE uses it as an orthogonal probe to introduce semantically heterogeneous reasoning signals into the teacher's orthogonal complement of its dominant subspace. This intervention steers the teacher toward more diverse reasoning trajectories and improves exploration beyond standard sampling. Experiments on mathematical benchmarks show that SOE improves average accuracy by 62.4\% and average sampling efficiency by 113.7\% over baseline methods, suggesting that geometric interventions can be effective for mitigating reasoning collapse in mathematical reasoning. We further provide preliminary evidence that SOE is also effective on logic and code generation benchmarks.

推理增强几何干预大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。