用语法特征评估子词切分是否符合形态规律,无需人工标注数据。
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
- 通过统计对齐方法将子词与形态特征匹配
- 在多种语言上相关性优于传统评测指标
- 适合无标注数据的低资源语言研究
我们提出一种新度量方法,用于评估子词分割在形态学上的合理性。不同于依赖人工标注的形态边界或召回率指标(这些在多数语言中难以获取或质量不一),本方法利用通用依存句法库(Universal Dependencies)或UniMorph等资源中的形态-句法特征。该度量通过IBM Model 1实现子词与形态特征的概率对齐。实验表明,该方法在不同形态系统语言间具有良好的可扩展性,且与传统形态边界召回率高度相关。
原文摘要 · Abstract (English)
We present a novel metric for the evaluation of the morphological plausibility of subword segmentation. Unlike the typically used morpheme boundary or retrieval F-score, which requires gold segmentation data that is either unavailable or of inconsistent quality across many languages, our approach utilizes morpho-syntactic features. These are available in resources such as Universal Dependencies or UniMorph for a much wider range of languages. The metric works by probabilistically aligning subwords with morphological features through an IBM Model 1. Our experiments show that the metric correlates well with traditional morpheme boundary recall while being more broadly applicable across languages with different morphological systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。