多语言混用数据提升大模型跨语言理解能力。
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

- 在四语环境下使用句级多语言混用数据进行指令微调。
- 在Belebele评测中,四种语言平均性能均获提升。
- 首次验证多语言混用比双语迁移更有效,适合多语种研究者。
近期研究表明,将多种语言混合在同一语境中的代码切换数据(CSD)可提升大语言模型(LLM)的跨语言迁移与多语言对齐能力。然而,现有研究主要聚焦英语与目标语言之间的双语迁移,缺乏对三语及以上多语言场景的探索。本文首次在英文、日文、韩文和中文四语环境下,开展多语言代码切换指令微调研究,并在Belebele评测集上评估多语言理解性能。实验表明,简单的句级多语言CSD能持续提升所有四种语言的平均表现,证明多语言代码切换在双语迁移之外仍具有效性。
原文摘要 · Abstract (English)
Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual transfer and multilingual alignment in large language models (LLMs). However, existing studies primarily focus on bilingual transfer between English and a target language, leaving multilingual settings involving three or more languages largely unexplored. In this work, we investigate multilingual code-switching instruction tuning across four languages: English, Japanese, Korean, and Chinese. We evaluate multilingual understanding on Belebele. Our experiments show that simple sentence-level multilingual CSD consistently improves average multilingual performance across all four languages, indicating that multilingual code-switching can be effective beyond bilingual transfer settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。