arXiv:2508.12662cs.CLcs.AI2025-08中稿 · NAACL被引 2

用合成混语数据微调,让低资源语言模型表现更公平。

Breaking Language Barriers: Equitable Performance in Multilingual Language Models

  • 用可控混语生成技术构造合成训练数据
  • 低资源语言推理准确率显著提升,高资源语言不变差
  • 适合关注多语言公平性与模型鲁棒性的研究者

当前大型语言模型在多语言交流与理解中表现强大,但在低资源语言(如印地语、斯瓦希里语)的常识推理任务中表现明显逊于高资源语言(如英语)。为实现不同语言群体间公平获取高质量模型输出,本文提出一种新方法:使用受控语言混杂技术生成合成混语文本,并用于微调语言模型。实验证明,该方法能显著提升低资源语言下的模型性能,同时保持或增强高资源语言的表现。此外,我们构建了一个基于CommonSenseQA的合成混语数据集,包含三种不同的语言比例配置。

原文摘要 · Abstract (English)

Cutting-edge LLMs have emerged as powerful tools for multilingual communication and understanding. However, LLMs perform worse in Common Sense Reasoning (CSR) tasks when prompted in low-resource languages (LRLs) like Hindi or Swahili compared to high-resource languages (HRLs) like English. Equalizing this inconsistent access to quality LLM outputs is crucial to ensure fairness for speakers of LRLs and across diverse linguistic communities. In this paper, we propose an approach to bridge this gap in LLM performance. Our approach involves fine-tuning an LLM on synthetic code-switched text generated using controlled language-mixing methods. We empirically demonstrate that fine-tuning LLMs on synthetic code-switched datasets leads to substantial improvements in LRL model performance while preserving or enhancing performance in HRLs. Additionally, we present a new dataset of synthetic code-switched text derived from the CommonSenseQA dataset, featuring three distinct language ratio configurations.

多语言公平性模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。