低资源语言在大模型中表现差,跨语言迁移能部分改善但效果因模型而异
Left Behind: Cross-Lingual Transfer as a Bridge for Low-Resource Languages in Large Language Models
- 用英文推理再翻译回低资源语言,提升部分双语模型表现
- 低资源语言准确率比英语低13.8至16.7个百分点
- 模型架构决定是否受益,无普适解法,适合关注多语言公平的开发者
我们通过在英语、哈萨克语和蒙古语中对八种大语言模型进行五种实验条件的基准测试,评估其在低资源语言上的表现。使用50个涵盖事实、推理、技术及文化背景的问题,评估2000条回复的准确性、流畅性和完整性。结果发现,英语与低资源语言之间存在13.8%至16.7%的一致性性能差距,模型虽保持表面流畅,但内容准确性显著下降。跨语言迁移——即让模型先以英文推理再翻译回目标语言——仅对双语架构带来+2.2至+4.3个百分点的增益,对英语主导型模型无效。研究证明当前大语言模型系统性忽视低资源语言群体,且有效缓解策略依赖于模型架构而非通用方法。
原文摘要 · Abstract (English)
We investigate how large language models perform on low-resource languages by benchmarking eight LLMs across five experimental conditions in English, Kazakh, and Mongolian. Using 50 hand-crafted questions spanning factual, reasoning, technical, and culturally grounded categories, we evaluate 2,000 responses on accuracy, fluency, and completeness. We find a consistent performance gap of 13.8-16.7 percentage points between English and low-resource language conditions, with models maintaining surface-level fluency while producing significantly less accurate content. Cross-lingual transfer-prompting models to reason in English before translating back-yields selective gains for bilingual architectures (+2.2pp to +4.3pp) but provides no benefit to English-dominant models. Our results demonstrate that current LLMs systematically underserve low-resource language communities, and that effective mitigation strategies are architecture-dependent rather than universal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。