arXiv:2505.21354cs.CLcs.LG2025-05被引 5

构建首个复杂孟加拉数学题数据集,用思维链提升大模型解题能力

Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning

  • 创建8792个带逐步解答的孟加拉数学题数据集,支持推理评估
  • 使用思维链提示使模型准确率达88%,显著优于常规提示
  • 通过低秩适配高效微调,适合资源有限的低语种研究

由于孟加拉语资源匮乏且需多步推理,解决孟加拉数学应用题仍是自然语言处理的重大挑战。现有模型在复杂题上表现不佳,主要因缺乏人工标注的孟加拉语数据集。为此,我们构建了SOMADHAN数据集,包含8792个复杂孟加拉数学题及其手动编写的分步解答,旨在支持语言代表性不足情境下的推理评估与模型开发。基于该数据集,我们评估了GPT-4o、GPT-3.5 Turbo、LLaMA系列、Deepseek和Qwen等大语言模型,在零样本与少样本提示下,结合或不结合思维链(CoT)推理。结果显示,思维链提示显著提升性能,尤其在需要多步逻辑的任务中。其中,LLaMA-3.3 70B在少样本思维链提示下达到88%最高准确率。我们还采用低秩适配(LoRA)实现高效微调,以极小计算成本使模型适应孟加拉数学题。本工作填补了孟加拉语NLP在推理任务上的关键空白,提供高质量数据集与可扩展解决方案,助力低资源语言公平研究与教育技术发展。

原文摘要 · Abstract (English)

Solving Bengali Math Word Problems (MWPs) remains a major challenge in natural language processing (NLP) due to the language's low-resource status and the multi-step reasoning required. Existing models struggle with complex Bengali MWPs, largely because no human-annotated Bengali dataset has previously addressed this task. This gap has limited progress in Bengali mathematical reasoning. To address this, we created SOMADHAN, a dataset of 8792 complex Bengali MWPs with manually written, step-by-step solutions. We designed this dataset to support reasoning-focused evaluation and model development in a linguistically underrepresented context. Using SOMADHAN, we evaluated a range of large language models (LLMs) - including GPT-4o, GPT-3.5 Turbo, LLaMA series models, Deepseek, and Qwen - through both zero-shot and few-shot prompting with and without Chain of Thought (CoT) reasoning. CoT prompting consistently improved performance over standard prompting, especially in tasks requiring multi-step logic. LLaMA-3.3 70B achieved the highest accuracy of 88% with few-shot CoT prompting. We also applied Low-Rank Adaptation (LoRA) to fine-tune models efficiently, enabling them to adapt to Bengali MWPs with minimal computational cost. Our work fills a critical gap in Bengali NLP by providing a high-quality reasoning dataset and a scalable framework for solving complex MWPs. We aim to advance equitable research in low-resource languages and enhance reasoning capabilities in educational and language technologies.

数学推理低资源语言思维链大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。