提出新指标评估跨语言迁移真实能力,发现小模型迁移没坏,大模型进步慢但整体在变好。
Are Multilingual Models Actually Improving? Isolating True Cross-Lingual Transfer

- 用硬度调整转移得分(HAT)分离源语言性能提升与真实跨语言迁移能力。
- 20个模型测试显示:模型越大,跨语言迁移提升越慢,但整体趋势向好。
- 适合研究多语言模型性能、评估跨语言泛化能力的学者与工程师参考。
跨语言迁移指模型将源语言中的能力泛化到资源较少的目标语言。现有评估方法混淆了源语言准确率提升与真实迁移能力的改善。本文提出一种新指标——硬度调整转移(HAT)得分,用于可靠衡量迁移强度,并基于此分析二十个多样化的语言模型和三个主流多语言基准的数据,得出三点结论:1)小模型的跨语言迁移并未失效;2)随着模型规模扩大,跨语言迁移的提升速度低于预期;3)总体而言,跨语言迁移能力随时间有明显进步。
原文摘要 · Abstract (English)
Cross-lingual transfer is a model's ability to generalize capabilities from well-represented source languages to under-represented target languages. Existing measures of a model's transfer strength conflate improvements in transfer with general improvements to accuracy in the source language. We advocate for an alternate metric that reliably captures transfer strength called Hardness Adjusted Transfer (HAT) Score, and use it to derive multiple insights on factors influencing transfer strength. Our analysis across twenty diverse language models and three popular mainstream multilingual benchmarks argues that 1) transfer in small models is not broken, 2) we are making slower than expected progress in cross-lingual transfer with model size, and 3) we have made clear progress over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。