arXiv:2608.30463cs.CL2026-08

首个多语言数学可解性检测基准,发现模型信念具有跨语言通用性。

More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs

论文配图:More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs
图 1 · 摘自论文原文
  • 构建中法希三语配对可解/不可解数学题数据集
  • 高资源语言数学推理更强但判断可信度更低
  • 模型对可解性的信念具有跨语言通用性

可解性检测是大语言模型进行数学推理中最具挑战的方面之一。尽管此前研究对此能力进行了广泛探讨,但分析主要局限于英语。因此,多语言失败究竟是源于内部可解性信念差异,还是语言表达障碍尚不明确。为填补这一空白,我们引入首个包含配对可解与不可解数学问题的多语言基准,将ReliableMath扩展至法语和希腊语。利用该基准,我们训练多语言探测器以预测可解性信念,并从行为、表征及可信度三个层面分析当前最先进大模型的可解性检测能力。结果表明,可解性信念作为一种基本且普遍存在的特征被编码,而更高资源语言(如英语)虽在数学推理表现上更优,但其可解性检测的可信度反而更低。

原文摘要 · Abstract (English)

Solvability detection is one of the most challenging aspects of mathematical reasoning for Large Language Models (LLMs). While prior work has studied this capability extensively, these analyses have been limited to English. Consequently, it remains unclear whether multilingual failures arise from differences in internal Solvability Belief or from language-dependent failures to express it. To address this gap, we introduce the first multilingual benchmark of paired solvable and unsolvable mathematical problems, extending ReliableMath to French and Greek. Using this, we train multilingual probes predicting Solvability Belief and analyze the solvability detection capabilities of state-of-the-art LLMs behaviorally, representationally, and in terms of faithfulness. We find that Solvability Belief is encoded as a largely universal, language-agnostic feature, and that higher-resource languages such as English, despite achieving stronger mathematical reasoning performance, exhibit lower solvability-detection faithfulness.

数学推理多语言可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。