比较两种文字统一方法在突厥语族中的跨语言迁移效果
Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages
- 用通用罗马化和突厥语专用文字统一方案处理11种突厥语
- 命名实体识别中两种方法表现相近,均显著优于单语基线
- 词性标注效果取决于语言特性和监督数据,无绝对优劣
密切相关的不同文字书写语言对多语言模型而言表面重叠极少,限制了跨语言迁移。本文比较了两种文字统一方法:通用的uroman罗马化与针对突厥语族的通用突厥文字(CTS)。我们在11种突厥语的音译维基百科语料上训练匹配的fastText模型,并在WikiANN命名实体识别和Universal Dependencies词性标注任务上评估。NER任务中,CTS与uroman表现无显著差异,两者均显著优于官方单语fastText基线。在POS任务中无统一优胜者:跨语言字符n-gram覆盖度由表示方式决定,而目标语言监督充足时,同语言覆盖度更为关键。尽管CANINE-c整体平均性能更高,但结构更简单的fastText系统在多个树库上仍具竞争力。总体而言,文字统一的有效性取决于语言、诱导的子词重叠度以及可用监督数据。
原文摘要 · Abstract (English)
Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare two approaches to script unification: the general-purpose uroman romanizer and the family-specific Common Turkic Script (CTS). We train matched fastText models on transliterated Wikipedia corpora from 11 Turkic languages and evaluate them on WikiANN named entity recognition and Universal Dependencies part-of-speech tagging. CTS and uroman show no significant difference on NER, while both substantially outperform the official monolingual fastText baselines. POS results reveal no universal winner: language-specific differences are associated with the cross-lingual character n-gram coverage induced by each representation, while within-language coverage becomes more important when target-language supervision is available. Although CANINE-c achieves higher overall POS averages, the substantially simpler fastText-based systems remain competitive on several treebanks. Overall, the effectiveness of script unification depends on the language, the induced subword overlap, and the available supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。