arXiv:2505.19848cs.CL2025-05被引 1

测试不同方法提升非洲语言数学推理能力

Improving Multilingual Math Reasoning for African Languages

  • 对比翻译与合成数据在低资源语言中的效果
  • 发现后训练阶段使用合成数据表现更优
  • 适合关注非洲语言AI研究的学者参考

低资源语言研究面临数据稀缺和计算资源受限的挑战。尽管大型语言模型(LLMs)主要在高资源语言上训练,但将其适配至非洲语言仍需专门技术。当前主流方法包括多阶段预训练与后训练范式,但最优策略尚不明确。本文系统评估了多种适应策略在扩展现有LLMs至非洲语言时的表现,聚焦数学推理任务,以Llama 3.1模型家族为基线。通过大量实验与消融分析,考察了翻译数据与合成数据、预训练与后训练阶段等组合的效果。结果表明,在后训练阶段使用合成数据能显著提升模型在非洲语言数学推理上的性能。

原文摘要 · Abstract (English)

Researchers working on low-resource languages face persistent challenges due to limited data availability and restricted access to computational resources. Although most large language models (LLMs) are predominantly trained in high-resource languages, adapting them to low-resource contexts, particularly African languages, requires specialized techniques. Several strategies have emerged for adapting models to low-resource languages in todays LLM landscape, defined by multi-stage pre-training and post-training paradigms. However, the most effective approaches remain uncertain. This work systematically investigates which adaptation strategies yield the best performance when extending existing LLMs to African languages. We conduct extensive experiments and ablation studies to evaluate different combinations of data types (translated versus synthetically generated), training stages (pre-training versus post-training), and other model adaptation configurations. Our experiments focuses on mathematical reasoning tasks, using the Llama 3.1 model family as our base model.

多语言低资源数学推理非洲语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。