跨语言数词系统结构标注与分析,助力比较研究
Annotating and Inferring Compositional Structures in Numeral Systems Across Languages
- 提出简洁有效的数词标注方案,支持计算机辅助编码
- 25种语言1-40的数词分析揭示底层与表层结构差异
- 发现异形变体是分词错误主因,子词分词不适用于低资源场景
全球语言中的数词系统在共时结构和历时演变上呈现丰富多样性。为实现跨语言数词系统的有效比较,需将其编码为标准化形式以对比基本属性。本文提出一种简单而高效的数词标注方案,并设计配套工作流程,实现计算机辅助编码,提供25种类型多样的语言从1到40的数词样本数据。我们对样本进行深入分析,重点比较底层与表层形态结构的系统性差异。进一步实验了自动化形态素分割模型,发现异形变体(allomorphy)是导致分割错误的主要原因。最后表明,子词分词算法在低资源场景下无法有效识别形态素。
原文摘要 · Abstract (English)
Numeral systems across the world's languages vary in fascinating ways, both regarding their synchronic structure and the diachronic processes that determined how they evolved in their current shape. For a proper comparison of numeral systems across different languages, however, it is important to code them in a standardized form that allows for the comparison of basic properties. Here, we present a simple but effective coding scheme for numeral annotation, along with a workflow that helps to code numeral systems in a computer-assisted manner, providing sample data for numerals from 1 to 40 in 25 typologically diverse languages. We perform a thorough analysis of the sample, focusing on the systematic comparison between the underlying and the surface morphological structure. We further experiment with automated models for morpheme segmentation, where we find allomorphy as the major reason for segmentation errors. Finally, we show that subword tokenization algorithms are not viable for discovering morphemes in low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。