利用语音文本对齐数据自动挖掘中文分词边界,提升跨领域分词效果
Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
- 通过语音文本强制对齐识别停顿作为候选分词边界
- 基于概率策略过滤不可靠边界,提升边界质量
- 提出完整训练策略,有效利用额外标注数据,适合跨域场景
受早期自然标注数据研究及语音文本融合处理启发,本文首次提出显式从语音-文本平行数据中挖掘分词边界。采用蒙特利尔强制对齐工具(MFA)对语音-文本数据进行字符级对齐,将停顿视为候选分词边界。通过对收集到的停顿进行详细分析,提出一种有效的基于概率的不可靠边界过滤策略。为更高效利用分词边界作为额外训练数据,进一步提出稳健的‘先完整后训练’(CTT)策略。在两个目标领域ZX和AISHELL2上开展跨领域中文分词实验,其中对AISHELL2人工标注约1000句作为评测数据。实验表明所提方法有效。
原文摘要 · Abstract (English)
Inspired by early research on exploring naturally annotated data for Chinese Word Segmentation (CWS), and also by recent research on integration of speech and text processing, this work for the first time proposes to explicitly mine word boundaries from speech-text parallel data. We employ the Montreal Forced Aligner (MFA) toolkit to perform character-level alignment on speech-text data, giving pauses as candidate word boundaries. Based on detailed analysis of collected pauses, we propose an effective probability-based strategy for filtering unreliable word boundaries. To more effectively utilize word boundaries as extra training data, we also propose a robust complete-then-train (CTT) strategy. We conduct cross-domain CWS experiments on two target domains, i.e., ZX and AISHELL2. We have annotated about 1,000 sentences as the evaluation data of AISHELL2. Experiments demonstrate the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。