arXiv:2512.05580cs.CL2025-12被引 2

用树状思维提升孟加拉语数学题解题准确率

Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems

  • 采用树状思维结构替代线性推理,避免错误传播
  • 在GPT-OSS-120B上达88%准确率,比线性推理高5个百分点
  • 适合低资源语言数学题求解,尤其对大模型有效

数学应用题是自然语言处理中最具挑战性的任务之一,因其需兼顾语言理解与多步数值推理。尽管链式思维(CoT)提示已展现潜力,但其线性结构常导致错误传播,限制整体效果。为此,我们系统研究了树状思维(ToT)在孟加拉语数学应用题中的应用,基于SOMADHAN数据集,评估了100个代表性问题在多个大语言模型(包括GPT-OSS和LLaMA系列)上的表现,对比标准提示、CoT与ToT策略。结果显示,CoT将基线准确率从78%(标准提示)提升至平均83%,而ToT进一步提升5个百分点,在GPT-OSS-120B上达到88%。该结果表明,ToT在中大型模型中尤为有效,对小模型提升有限。整体而言,本研究确立了ToT作为低资源语言如孟加拉语数学题求解的稳健框架,并证明结构化推理方法能提供更可靠、全局一致的解题结果,为多语言NLP中的推理策略发展提供新方向。

原文摘要 · Abstract (English)

Mathematical Word Problems (MWPs) are among the most challenging tasks in natural language processing because they require both linguistic understanding and multi-step numerical reasoning. While Chain-of-Thought (CoT) prompting has shown promise, its linear structure often propagates errors, limiting overall effectiveness. To address this limitation, we present the a systematic study of Tree-of-Thought (ToT) reasoning for Bengali MWPs using the SOMADHAN dataset. Owing to computational and token-cost constraints, we evaluate a curated set of 100 representative problems across multiple large language models (LLMs), including GPT-OSS and LLaMA variants, under standard prompting, CoT, and ToT strategies. Our results show that CoT improves baseline accuracy from 78% (standard prompting) to 83% on average, while ToT further increases performance by up to 5 percentage points, achieving 88% accuracy with GPT-OSS-120B. These improvements highlight that ToT is particularly effective in medium-to-large-scale models but may offer less advantage for smaller ones. Overall, our findings establish ToT as a robust framework for solving mathematical problems in low-resource languages such as Bengali. More broadly, this study shows that structured reasoning methods like ToT can provide more reliable and globally consistent outcomes than CoT, paving the way for better reasoning strategies in multilingual NLP.

数学推理树状思维低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。