用语料数据补全语言类型学特征,提升多语言模型适应性。
data2lang2vec: Data Driven Typological Features Completion
- 基于语料和词性标注,预测缺失的语言类型学特征。
- 在1749种语言上实现超70%准确率,覆盖率达28.9%以上。
- 新评估方式更贴近真实缺失场景,适合多语言研究者。
语言类型学数据库通过增强模型对多样化语言结构的适应性,推动多语言自然语言处理发展。广泛使用的 lang2vec 工具包整合了多个此类数据库,但其覆盖率仍仅为28.9%。以往工作通常基于其他语言特征预测缺失值,或仅针对单一特征进行建模;本文提出利用文本数据实现更精准的特征预测。为此,我们构建了一个多语言词性标注器(POS tagger),在1749种语言上达到超过70%的准确率,并结合外部统计特征与多种机器学习算法进行实验。此外,我们引入更贴近实际应用的评估设置,聚焦于最可能缺失的类型学特征,结果表明该方法在两种评估环境下均优于先前工作。
原文摘要 · Abstract (English)
Language typology databases enhance multi-lingual Natural Language Processing (NLP) by improving model adaptability to diverse linguistic structures. The widely-used lang2vec toolkit integrates several such databases, but its coverage remains limited at 28.9\%. Previous work on automatically increasing coverage predicts missing values based on features from other languages or focuses on single features, we propose to use textual data for better-informed feature prediction. To this end, we introduce a multi-lingual Part-of-Speech (POS) tagger, achieving over 70\% accuracy across 1,749 languages, and experiment with external statistical features and a variety of machine learning algorithms. We also introduce a more realistic evaluation setup, focusing on likely to be missing typology features, and show that our approach outperforms previous work in both setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。