arXiv:2512.06169cs.CL2025-12

为非连续构词语言设计新型分词器,提升语音识别标注效率。

Morphologically-Informed Tokenizers for Languages with Non-Concatenative Morphology: A case study of Yoloxóchtil Mixtec ASR

  • 提出两种非线性分词方案,保留声调构词信息。
  • 分段与声调分词器在词错误率上优于传统模型。
  • 适合低资源语言语音标注,尤其混合语系研究者。

本文研究使用构词信息分词器辅助音视频语料库的逐行注释,以提高效率并减轻人工标注负担。针对非连续构词的约洛克奇特尔米斯特克语(Yoloxóchtil Mixtec, YM),提出两种新分词方法:一种是仅提取声调的分段与声调分词器;另一种是预测词切分的流程序列分词器,可实现端到端语音识别中同时输出分段与未分段转录。实验表明,这些新分词器在词错误率(WER)上表现优于BPE和Unigram模型,但字符错误率(CER)未达同等水平。通过构词与信息论指标分析,发现其与下游性能存在预测相关性。结果表明,专为非连续构词语言设计的非线性分词器,在语音识别任务中具有竞争力,未来需进一步验证其在下游任务中的适用性。

原文摘要 · Abstract (English)

This paper investigates the impact of using morphologically-informed tokenizers to aid and streamline the interlinear gloss annotation of an audio corpus of Yoloxóchitl Mixtec (YM) using a combination of ASR and text-based sequence-to-sequence tools, with the goal of improving efficiency while reducing the workload of a human annotator. We present two novel tokenization schemes that separate words in a nonlinear manner, preserving information about tonal morphology as much as possible. One of these approaches, a Segment and Melody tokenizer, simply extracts the tones without predicting segmentation. The other, a Sequence of Processes tokenizer, predicts segmentation for the words, which could allow an end-to-end ASR system to produce segmented and unsegmented transcriptions in a single pass. We find that these novel tokenizers are competitive with BPE and Unigram models, and the Segment-and-Melody model outperforms traditional tokenizers in terms of word error rate but does not reach the same character error rate. In addition, we analyze tokenizers on morphological and information-theoretic metrics to find predictive correlations with downstream performance. Our results suggest that nonlinear tokenizers designed specifically for the non-concatenative morphology of a language are competitive with conventional BPE and Unigram models for ASR. Further research will be necessary to determine the applicability of these tokenizers in downstream processing tasks.

语音识别分词器低资源语言构词分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。