arXiv:2510.01831cs.CL2025-10中稿 · MathNLP 2025: The …被引 3

LLM解数学题常因句式陌生出错,改写句式后正确率显著提升

Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors

  • 通过改写问题句式,保留语义但降低结构复杂度,提升模型正确率
  • 句式复杂度越高(DLT得分高),错误率越明显,跨数据集一致
  • 揭示了语法形式与模型推理的脆弱关联,适合研究模型泛化与可解释性的人参考

大型语言模型在数学问题求解上表现出色,但对训练分布外的句式变化常出现错误。我们识别出一种系统性失败模式——句法盲区:模型将熟悉的推理策略误用于语义清晰但句式陌生的问题。这些错误并非源于数学能力不足,而是表面形式与内部表征之间的脆弱耦合所致。为验证此假设,我们采用正确例题中的句法模板重写错误解答的问题,保持语义不变的同时降低结构复杂度,结果多数情况获得正确答案。我们使用基于依存局部性理论(DLT)的指标量化句法复杂度,发现更高DLT得分对应更高的错误率,且该现象在多个数据集上均成立。研究提示,许多推理错误源于结构不匹配而非概念难度,而语法敏感干预可揭示并缓解这类归纳失败。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate strong mathematical problem-solving abilities but frequently fail on problems that deviate syntactically from their training distribution. We identify a systematic failure mode, syntactic blind spots, in which models misapply familiar reasoning strategies to problems that are semantically straightforward but phrased in unfamiliar ways. These errors are not due to gaps in mathematical competence, but rather reflect a brittle coupling between surface form and internal representation. To test this, we rephrase incorrectly answered questions using syntactic templates drawn from correct examples. These rephrasings, which preserve semantics while reducing structural complexity, often lead to correct answers. We quantify syntactic complexity using a metric based on Dependency Locality Theory (DLT), and show that higher DLT scores are associated with increased failure rates across multiple datasets. Our findings suggest that many reasoning errors stem from structural misalignment rather than conceptual difficulty, and that syntax-aware interventions can reveal and mitigate these inductive failures.

大模型数学推理句法盲区可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。