arXiv:2509.20129cs.CL2025-09中稿 · EMNLP

压缩语言类型学特征,提升跨语言模型效果

Less is More: The Effectiveness of Compact Typological Language Representations

  • 通过特征筛选与填补优化高维稀疏的类型学特征
  • 小规模特征集使语言距离度量更准确,性能提升明显
  • 适合低资源语言的多语言NLP任务研究者

语言特征数据集如URIEL+对建模跨语言关系很有价值,但其高维度和稀疏性(尤其在低资源语言中)限制了距离度量的效果。本文提出一个流水线方法,结合特征选择与缺失值填补,优化URIEL+类型学特征空间,生成紧凑且可解释的类型学表示。在语言距离对齐及下游任务上的评估表明,缩减后的类型学表示能产生更有效的距离度量,并提升多语言自然语言处理应用的表现。

原文摘要 · Abstract (English)

Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance metrics. We propose a pipeline to optimize the URIEL+ typological feature space by combining feature selection and imputation, producing compact yet interpretable typological representations. We evaluate these feature subsets on linguistic distance alignment and downstream tasks, demonstrating that reduced-size representations of language typology can yield more informative distance metrics and improve performance in multilingual NLP applications.

语言学特征优化多语言NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。