arXiv:2608.08503cs.AIcs.CL2026-08

对比了孟加拉语数学推理中答案与思维链监督的效果。

MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models

论文配图:MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
图 1 · 摘自论文原文
  • 用相同数据和训练条件,对比答案仅监督与思维链监督
  • 强模型上思维链无显著提升,弱模型提升18.56分
  • 主要优势是语言一致性与推理过程可解释性

数学推理在低资源语言如孟加拉语中仍具挑战。本文研究教师生成的孟加拉语思维链(CoT)监督是否优于普通监督微调。构建了由GPT-5.4生成推理过程的 extsc{MathShikkha}数据集,对四个4B–7B规模的学生模型进行匹配协议下的微调,仅训练目标不同:答案仅监督与思维链监督。在领域内,三个较强模型的思维链未带来显著提升(配对自举95%置信区间包含零;精确McNemar p ≥ 0.17),尽管生成了15–52×更多词元;但对较弱的4B模型显著提升18.56分(p < 0.0001)。在更大的、经过污染审计的BanglaMATH基准上,该模式反转:所有四模型中思维链均显著优于答案监督,提升20.1–28.1分(均p < 0.0001)。答案监督导致三个模型域外准确率低于基线,而思维链对全部四模型保持或提升表现。人工评估显示,思维链在推理内容上未显著优于基线(κ=0.76–1.00),但显著提升目标语言一致性与推理可检视性。总体而言,该监督方式的价值取决于模型能力与分布偏移:其核心优势在于语言适配性、可审计性与域外鲁棒性,而非提升域内推理有效性。

原文摘要 · Abstract (English)

Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B--7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95\% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15--52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p < 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1--28.1 points (all $p < 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen's $κ= 0.76$--$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision's value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.

数学推理低资源语言思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。