arXiv:2605.27715cs.CL2026-05

用图结构诊断多语言数学推理缺陷,提升低资源语言表现

Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs

论文配图:Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
图 1 · 摘自论文原文
  • 构建有向无环追踪图(DATG),将推理过程映射为语言无关的数学节点与依赖
  • 发现非英语推理中数学节点覆盖不足、依赖关系错误,尤其在低资源语言更严重
  • 提出两种测试时控制方法,有效改善低资源语言的数学推理准确率

大型推理模型在英语数学题上表现优异,但在低中资源语言中可靠性显著下降。传统观点认为这是由于理解问题表述能力不足,但我们发现:即使问题用英语给出,仅改变推理语言也会大幅降低准确率,说明语言影响推理执行本身。为此,我们提出定向无环追踪图(DATG)框架,将推理轨迹映射为语言无关的数学锚点与依赖关系,可对齐目标语言轨迹与参考图谱,评估其是否覆盖必要数学节点、遵守依赖边、避免有害操作。在Qwen3系列跨12种语言的实验表明,非英语推理普遍存在锚点覆盖率低、依赖保真度差的问题,尤其在低资源语言中更为明显。基于此诊断,我们设计了Loop-Retry与Formula-Retry两种测试时控制策略,针对暴露的失效模式,显著提升了低资源语言的推理性能。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) achieve strong mathematical reasoning performance in English, but remain much less reliable in many low- and medium-resource languages. This gap is often explained as a failure to understand non-English problem statements. We show that this view is incomplete: even when the problem is given in English, controlling the model's reasoning language can substantially reduce accuracy, suggesting that language also affects reasoning execution itself. To study this effect, we introduce DATG, a Directed Acyclic Trace Graph framework that maps reasoning traces to language-independent mathematical anchors and dependencies. This allows us to align target-language traces with reference DAGs and measure whether they cover required mathematical nodes, respect dependency edges, and avoid harmful mathematical actions. Experiments on the Qwen3 series across 12 languages show that non-English reasoning often suffers from reduced anchor coverage and weaker dependency fidelity, especially in low-resource languages. Motivated by this diagnosis, we propose Loop-Retry and Formula-Retry, two simple test-time controls targeting DATG-exposed failure modes, and show that they consistently improve target-language reasoning performance in low-resource languages.

数学推理多语言推理诊断大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。