arXiv:2510.27183cs.CL2025-10中稿 · LREC 2026被引 1

扩充语言数据集,提升低资源语言支持能力

Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

  • 新增7488种语言的书写系统向量,增强语言表征
  • 通过整合Glottolog扩展至超2万语言,覆盖提升10倍以上
  • 跨语言迁移任务中对低资源语言性能最高提升6%

URIEL+ 语言知识库通过地理、谱系和类型学向量支持多语言研究,但存在数据稀疏问题(如特征缺失、语言条目不全、谱系覆盖有限),限制其在跨语言迁移中的应用,尤其对低资源语言支持不足。为缓解此问题,本文扩展URIEL+:引入7,488种语言的书写系统向量,集成Glottolog增加18,710个语言条目,并通过谱系传播类型学与书写系统特征,实现对26,449种语言的谱系推断扩展。改进后,书写系统向量的特征稀疏率降低14%,语言覆盖增加最多19,015种(达1,007%),推断质量指标最高提升35%。在面向低资源语言的跨语言迁移任务基准测试中,性能相比原始URIEL+出现部分分歧,但在某些设置下最高提升6%。

原文摘要 · Abstract (English)

The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.

语言模型多语言数据扩充

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。