arXiv:2503.19979cs.CL2025-03中稿 · NAACL被引 10

找对翻译语言关键在类型学与数据特征,不靠盲目试错

Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging

  • 用类型学和数据特征联合判断最佳翻译语言
  • 词重叠率、词频比、语系距离是核心影响因素
  • 适合研究跨语言标注与模型迁移的学者参考

跨语言迁移学习是缓解数据稀缺的重要手段,但选择合适的迁移语言仍具挑战。本文系统考察语言类型学、训练数据和模型架构在跨语言词性标注迁移中的作用。我们分析了数据集特异性与细粒度类型学特征对迁移语言排序的影响,使用两种不同来源的形态句法特征。不同于以往基于双语biLSTM的研究,本工作扩展至更现代的预训练多语言模型零样本预测流程。通过构建一系列迁移语言排序系统,评估不同特征输入在各类架构下的表现。结果显示,词重叠率、词频比(type-token ratio)和语系距离在所有架构中均位列前茅。研究发现,类型学与数据依赖特征的结合能获得最优排序效果,且任一特征组单独使用亦可实现良好性能。

原文摘要 · Abstract (English)

Cross-lingual transfer learning is an invaluable tool for overcoming data scarcity, yet selecting a suitable transfer language remains a challenge. The precise roles of linguistic typology, training data, and model architecture in transfer language choice are not fully understood. We take a holistic approach, examining how both dataset-specific and fine-grained typological features influence transfer language selection for part-of-speech tagging, considering two different sources for morphosyntactic features. While previous work examines these dynamics in the context of bilingual biLSTMS, we extend our analysis to a more modern transfer learning pipeline: zero-shot prediction with pretrained multilingual models. We train a series of transfer language ranking systems and examine how different feature inputs influence ranker performance across architectures. Word overlap, type-token ratio, and genealogical distance emerge as top features across all architectures. Our findings reveal that a combination of typological and dataset-dependent features leads to the best rankings, and that good performance can be obtained with either feature group on its own.

跨语言词性标注迁移学习类型学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。