跨语言数字谜题暴露大模型数学推理短板
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
- 用多语言数字谜题分离语言与数学因素
- 显式符号标记才能让模型正确解题
- 适合研究模型推理机制的学者
不同语言的数字符号系统差异巨大,人类能自如应对,但大语言模型在涉及跨语言数字符号系统的语言-数学谜题上表现不佳。我们通过一系列实验探究其原因,发现只有当问题中的数学运算明确使用已知符号(如“+”、“×”)标注时,模型才能稳定解题。消融实验进一步表明,数字符号的构造与组合规则对性能影响显著。人类能利用语言理解推断数字的隐含结构,而大模型缺乏这种对隐含构式结构的感知能力。结论是:从人类尺度数据中灵活推断隐含规则,仍是当前推理模型的核心挑战。
原文摘要 · Abstract (English)
Across languages, numeral systems vary widely in how they construct and combine numbers. While humans consistently learn to navigate this diversity, large language models (LLMs) struggle with linguistic-mathematical puzzles involving cross-linguistic numeral systems, which humans can learn to solve successfully. We investigate why this task is difficult for LLMs through a series of experiments that untangle the linguistic and mathematical aspects of numbers in language. Our experiments establish that models cannot consistently solve such problems unless the mathematical operations in the problems are explicitly marked using known symbols ($+$, $\times$, etc., as in "twenty + three"). In further ablation studies, we probe how individual parameters of numeral construction and combination affect performance. While humans use their linguistic understanding of numbers to make inferences about the implicit compositional structure of numerals, LLMs seem to lack this notion of implicit numeral structure. We conclude that the ability to flexibly infer compositional rules from implicit patterns in human-scale data remains an open challenge for current reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。