用语音辅助文本恢复阿拉伯语重音符号,更快更准。
Constrained CTC Decoding for Efficient Diacritic Restoration
- 基于CTC的非自回归方法,解码时加入字符级约束
- 在ArVoice和ClArTTS上误差率显著降低
- 适合需要高效准确恢复重音符号的研究者
本文研究阿拉伯语语音转写中的重音符号恢复问题。多数语音数据缺乏重音符号,限制了对细微语音差异的建模。近年来,语音模态被用于补充基于文本的重音恢复方法。我们提出一种基于连接时序分类(CTC)的高效非自回归语音到文本重音恢复方法。通过从无重音转写构建字符级重音化词网,并在解码中强制限制为合法重音实现,提升准确性。在古典阿拉伯语和现代标准阿拉伯语测试集(ArVoice 和 ClArTTS)上,相比计算复杂的多模态基线,本方法在两个数据集上均实现统计显著的重音错误率下降,证明其兼具性能与效率优势。
原文摘要 · Abstract (English)
In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been explored as a way to complement text-based diacritic restoration efforts. We propose an efficient non-autoregressive approach for speech-to-text diacritization based on Connectionist Temporal Classification (CTC). Our method incorporates hard constraints during decoding by constructing a character-level diacritization lattice from an undiacritized transcript and restricting hypotheses to valid diacritized realizations. We evaluate on Classical Arabic and Modern Standard Arabic test sets (namely, ArVoice and ClArTTS) against a more computationally-complex multi-modal diacritic restoration baseline, and show statistically significant reductions in diacritic error rates in both, demonstrating that the proposed approach offers both performance and efficiency gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。