用字典增强解码,自动标注日语语音音素与语调标签。
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
- 基于真值文本微调ASR模型,同步输出音节与标注标签。
- 利用词典先验纠正音素标注错误,提升准确性。
- 自动生成标签可媲美人工标注,适合构建高质量日语语音合成数据集。
本文提出一种对给定音频-文本对进行音素和语调标注的方法,旨在构建日语文本到语音(TTS)数据集。方法通过在真实文本条件下微调大规模预训练自动语音识别(ASR)模型,实现音节级音素与标注标签的联合输出。为进一步修正音素标注错误,采用融合词典先验知识的解码策略。客观评估结果表明,该方法优于仅依赖文本或音频的现有方法。主观评估显示,使用本方法标注标签训练的TTS模型所生成语音的自然度,与使用人工标注训练的模型相当。
原文摘要 · Abstract (English)
In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained automatic speech recognition (ASR) model, conditioned on ground truth transcripts, to simultaneously output phrase-level graphemes and annotation labels. To further correct errors in phonemic labeling, we employ a decoding strategy that utilizes dictionary prior knowledge. The objective evaluation results demonstrate that our proposed method outperforms previous approaches relying solely on text or audio. The subjective evaluation results indicate that the naturalness of speech synthesized by the TTS model, trained with labels annotated using our method, is comparable to that of a model trained with manual annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。