arXiv:2603.27651cs.CL2026-03

针对非洲语言的跨语言迁移,智能分配标注预算提升效果

Budget-Xfer: Budget-Constrained Source Language Selection for Cross-Lingual Transfer to African Languages

  • 将多源语言选择建模为预算约束下的资源分配问题
  • 多源迁移显著优于单源,尤其在数据利用不充分时
  • 嵌入相似性并非万能,任务不同效果差异明显

跨语言迁移学习通过利用高资源语言的标注数据,使低资源语言的自然语言处理成为可能。然而,现有源语言选择策略比较未控制总训练数据量,导致语言选择效应与数据量效应混淆。我们提出 Budget-Xfer 框架,将多源跨语言迁移建模为预算约束下的资源分配问题。给定固定标注预算 B,该框架联合优化应包含哪些源语言以及从每种语言分配多少数据。我们在三个非洲目标语言(豪萨语、约鲁巴语、斯瓦希里语)上,对命名实体识别和情感分析任务,使用两种多语言模型进行了 288 组实验。结果表明:(1) 多源迁移显著优于单源迁移(Cohen's d = 0.80 到 1.98),主因是结构化的预算未充分利用瓶颈;(2) 多源策略间差异微小且不显著;(3) 嵌入相似性作为选择代理的价值取决于任务,随机选择在命名实体识别中表现优于基于相似性的选择,但在情感分析中则不然。

原文摘要 · Abstract (English)

Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data from higher-resource sources, yet existing comparisons of source language selection strategies do not control for total training data, confounding language selection effects with data quantity effects. We introduce Budget-Xfer, a framework that formulates multi-source cross-lingual transfer as a budget-constrained resource allocation problem. Given a fixed annotation budget B, our framework jointly optimizes which source languages to include and how much data to allocate from each. We evaluate four allocation strategies across named entity recognition and sentiment analysis for three African target languages (Hausa, Yoruba, Swahili) using two multilingual models, conducting 288 experiments. Our results show that (1) multi-source transfer significantly outperforms single-source transfer (Cohen's d = 0.80 to 1.98), driven by a structural budget underutilization bottleneck; (2) among multi-source strategies, differences are modest and non-significant; and (3) the value of embedding similarity as a selection proxy is task-dependent, with random selection outperforming similarity-based selection for NER but not sentiment analysis.

跨语言迁移非洲语言预算分配NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。