用英语做桥梁,让大模型更好解韩语数学题。
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
- 以英语为锚点,分三步理解、求解、翻译跨语言数学题。
- 在韩语数学题集上提升10.91%,差距从11.6%缩小至0.7%。
- 方法可泛化到不同韩语领域,适合多语言推理研究者。
大型语言模型在复杂推理任务中表现优异,但在低资源语言(如韩语)上仍存在显著性能差距。为探究这一现象,我们构建了包含8,011个英韩双语数学题的基准数据集HRM8K。系统分析发现,性能差异主要源于对非英语输入的理解困难,而非推理能力不足。为此,我们提出UST(理解、求解、翻译)方法,将英语作为推理和求解的锚点。通过在13万条合成数据上微调,该方法在HRM8K上实现10.91%的性能提升,使多语言差距从11.6%降至0.7%。此外,成果在不同韩语领域具有良好泛化性,表明机器可验证内容训练出的能力可迁移至其他场景。相关数据集、训练数据与模型均已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate exceptional performance on complex reasoning tasks. However, despite their strong reasoning capabilities in high-resource languages (e.g., English and Chinese), a significant performance gap persists in other languages. To investigate this gap in Korean, we introduce HRM8K, a benchmark comprising 8,011 English-Korean parallel bilingual math problems. Through systematic analysis of model behaviors, we identify a key finding: these performance disparities stem primarily from difficulties in comprehending non-English inputs, rather than limitations in reasoning capabilities. Based on these findings, we propose UST (Understand, Solve, and Translate), a method that strategically uses English as an anchor for reasoning and solution generation. By fine-tuning the model on 130k synthetically generated data points, UST achieves a 10.91% improvement on the HRM8K benchmark and reduces the multilingual performance gap from 11.6% to 0.7%. Additionally, we show that improvements from UST generalize effectively to different Korean domains, demonstrating that capabilities acquired from machine-verifiable content can be generalized to other areas. We publicly release the benchmark, training dataset, and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。