用两阶段微调提升低资源语言的推理能力
LinguaLIFT: An Effective Two-stage Instruction Tuning Framework for Low-Resource Language Reasoning
- 通过代码混杂方式实现跨语言对齐,仅用英文数据迁移推理能力
- 在21种低资源语言上显著超越基线模型
- 适合关注多语言公平性与低资源语言研究的读者
大语言模型虽具备出色的多语言推理能力,但高、低资源语言间仍存在性能差距,源于预训练语料的语言不平衡,且现有推理基准缺乏低资源语言覆盖,加剧了评估偏差。为此,我们提出LinguaLIFT——一种面向低资源语言推理的两阶段指令微调框架。该框架采用语言对齐层,在无需多语言指令或平行语料的情况下,以代码混杂方式捕捉多语言对齐,通过仅含英文的指令微调数据将跨语言推理能力迁移至低资源语言。为全面评估多语言推理能力,我们构建了涵盖21种低资源、17种中等资源和10种高资源语言的多语言数学世界问题(MMWP)基准。实验表明,LinguaLIFT在MMWP及四个主流基准上均优于多个竞争基线。
原文摘要 · Abstract (English)
Large language models (LLMs) have exhibited impressive multilingual reasoning capabilities, driven by extensive multilingual pre-training corpora and instruction fine-tuning data. However, a performance gap exists between high- and low-resource language reasoning tasks due to the language imbalance in the pre-training corpus, which is exacerbated by evaluation bias in existing reasoning benchmarks lacking low-resource language coverage. To alleviate this issue, we propose LinguaLIFT, a two-stage instruction tuning framework for advancing low-resource language reasoning. LinguaLIFT employs a language alignment layer to capture multilingual alignment in a code-switched tuning way without requiring multilingual instruction or parallel data, thereby transferring the cross-lingual reasoning capabilities to low-resource languages through English-only instruction tuning data. To comprehensively evaluate the multilingual reasoning capabilities, we introduce the Multilingual Math World Problem (MMWP) benchmark, which spans 21 low-resource, 17 medium-resource, and 10 high-resource languages. Experimental results show that LinguaLIFT outperforms several competitive baselines across MMWP and four widely used benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。