精选语言子集可让低资源语言翻译更准且更省数据
Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages
- 用语言多样性高的小样本替代全语言对齐,提升效率
- 实验显示选对语言子集,低资源语言效果反而更好
- 适合资源少、想低成本提升多语言模型的团队
对齐是提升多语言模型跨语言迁移的有效策略,但实证结果常不一致,尤其在与英语语系差异大或资源匮乏的语言上表现不稳定。现有词对齐工具通常依赖高质量平行语料,而这对许多低资源语言(LRLs)难以获取。本文通过系统实验研究:是否必须使用所有语言进行对齐?结果表明,精心挑选的、语言多样化的子集能实现与全量对齐相当甚至更优的跨语言迁移效果,尤其对未见的低资源语言。这说明有效对齐无需覆盖全部语言,可显著降低数据收集成本,同时保持高效与鲁棒性。
原文摘要 · Abstract (English)
Realignment is a promising strategy to improve cross-lingual transfer in multilingual language models. However, empirical results are mixed and often unreliable, particularly for typologically distant or low-resource languages (LRLs) compared to English. Moreover, word realignment tools often rely on high-quality parallel data, which can be scarce or noisy for many LRLs. In this work, we conduct an extensive empirical study to investigate whether realignment truly benefits from using all available languages, or if strategically selected subsets can offer comparable or even improved cross-lingual transfer, and study the impact on LRLs. Our controlled experiments show that realignment can be particularly effective for LRLs and that using carefully selected, linguistically diverse subsets can match full multilingual alignment, and even outperform it for unseen LRLs. This indicates that effective realignment does not require exhaustive language coverage and can reduce data collection overhead, while remaining both efficient and robust when guided by informed language selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。