多语言训练未必提升零样本迁移,关键在数据与评估差异。
Multilinguality Does not Make Sense: Investigating Factors Behind Zero-Shot Transfer in Sense-Aware Tasks
- 对比28种语言,发现多语言训练非零样本迁移核心因素。
- 预训练与微调数据差异、评测偏差更影响跨语言表现。
- 适合关注低资源语言与跨语言迁移机制的研究者。
跨语言迁移是现代自然语言处理的核心,使模型能在未训练过的语言上执行任务。普遍认为,训练语言越多,零样本迁移效果越好。我们针对义项消歧和词汇语义演变两类感知语义的任务进行检验,发现多语言训练并非有效迁移的必要条件。大规模分析覆盖28种语言表明,预训练与微调数据差异及评测偏差更能解释多语言带来的表观优势。我们还发布了微调模型并提供实证基线以支持未来研究。尽管聚焦两类语义感知任务,研究结果对跨语言迁移具有广泛启示,尤其适用于低资源语言场景。
原文摘要 · Abstract (English)
Cross-lingual transfer is central to modern NLP, enabling models to perform tasks in languages different from those they were trained on. A common assumption is that training on more languages improves zero-shot transfer. We test this on sense-aware tasks-polysemy and lexical semantic change-and find that multilinguality is not necessary for effective transfer. Our large-scale analysis across 28 languages reveals that other factors, such as differences in pretraining and fine-tuning data and evaluation artifacts, better explain the perceived benefits of multilinguality. We also release fine-tuned models and provide empirical baselines to support future research. While focused on two sense-aware tasks, our findings offer broader insights into cross-lingual transfer, especially for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。