用嵌入相似性预测非洲语言跨语言迁移效果,找到有效指标。
Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages
- 测试五种相似性度量,发现余弦差和检索指标能可靠预测迁移效果。
- 相关系数达0.4-0.6,而CKA几乎无预测力(ρ≈0.1)。
- 不同模型结果相反,需按模型单独验证,适合低资源语言研究者。
跨语言迁移对构建低资源非洲语言的NLP系统至关重要,但从业者缺乏可靠的源语言选择方法。本文系统评估了五种嵌入相似性度量,在涵盖3个NLP任务、3个非洲语系多语言模型、12种来自四大语系语言的816次迁移实验中进行检验。结果表明,余弦差和基于检索的度量(P@1、CSLS)能可靠预测迁移成功(ρ=0.4–0.6),而CKA的预测能力可忽略不计(ρ≈0.1)。关键发现:跨模型合并数据时相关性符号反转(辛普森悖论),因此必须针对每个模型单独验证。嵌入相似性度量的预测能力与URIEL语言类型学相当。研究为源语言选择提供实证指导,并强调模型特异性分析的重要性。
原文摘要 · Abstract (English)
Cross-lingual transfer is essential for building NLP systems for low-resource African languages, but practitioners lack reliable methods for selecting source languages. We systematically evaluate five embedding similarity metrics across 816 transfer experiments spanning three NLP tasks, three African-centric multilingual models, and 12 languages from four language families. We find that cosine gap and retrieval-based metrics (P@1, CSLS) reliably predict transfer success ($ρ= 0.4-0.6$), while CKA shows negligible predictive power ($ρ\approx 0.1$). Critically, correlation signs reverse when pooling across models (Simpson's Paradox), so practitioners must validate per-model. Embedding metrics achieve comparable predictive power to URIEL linguistic typology. Our results provide concrete guidance for source language selection and highlight the importance of model-specific analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。