用无监督方法检测科克博拉克语词表中的发音异常,提升语言记录数据质量。
Automated Quality Control for Language Documentation: Detecting Phonotactic Inconsistencies in a Kokborok Wordlist
- 基于音节和字符级发音规则,自动识别词表中异常条目。
- 音节特征显著优于字符级别基线,召回率较高。
- 适合资源匮乏语言的田野工作者用于数据质检。
语言记录中的词汇数据常包含拼写错误和未标注的借词,可能误导语言分析。本文提出无监督异常检测方法,用于识别多语种科克博拉克语词表(与孟加拉语混合)中的发音不一致现象。通过字符级与音节级发音特征,算法可定位潜在拼写错误与借词。尽管精确率与召回率因异常细微而有限,但音节感知特征显著优于字符级基线。高召回策略为田野工作者提供系统化标记机制,助力低资源语言记录的数据质量提升。
原文摘要 · Abstract (English)
Lexical data collection in language documentation often contains transcription errors and undocumented borrowings that can mislead linguistic analysis. We present unsupervised anomaly detection methods to identify phonotactic inconsistencies in wordlists, applying them to a multilingual dataset of Kokborok varieties with Bangla. Using character-level and syllable-level phonotactic features, our algorithms identify potential transcription errors and borrowings. While precision and recall remain modest due to the subtle nature of these anomalies, syllable-aware features significantly outperform character-level baselines. The high-recall approach provides fieldworkers with a systematic method to flag entries requiring verification, supporting data quality improvement in low-resourced language documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。