arXiv:2506.04527cs.SDcs.CL2025-06中稿 · INTERSPEECH 2025被引 2

让语音标签与拼音保持一致,提升语音合成和口音识别效果

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

  • 用BERT特征隐式加拼音条件,推理时显式剔除不一致标签
  • 使拼音与音素/语调标注一致性显著提升,口音估计准确率提高
  • 适合需要拼音对齐的语音合成、口音分析等任务

我们提出一种模型,生成与拼音一致的音素和语调标签。不同于以往仅微调预训练语音识别模型的方法,该模型通过两种方式将标签生成与对应拼音关联:1)利用预训练BERT特征的提示编码器实现隐式拼音条件;2)在推理阶段显式剔除与拼音不一致的标签候选。该方法可构建语音、标签与拼音的并行数据,适用于文本到语音合成、从文本估算口音等多种下游任务。实验表明,该方法显著提升了拼音与预测标签的一致性;在口音估计任务中进一步验证,所生成的并行数据有效提升了估计精度。

原文摘要 · Abstract (English)

We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation on corresponding graphemes by two methods: 1) Add implicit grapheme conditioning through prompt encoder using pre-trained BERT features. 2) Explicitly prune the label hypotheses inconsistent with the grapheme during inference. These methods enable obtaining parallel data of speech, the labels, and graphemes, which is applicable to various downstream tasks such as text-to-speech and accent estimation from text. Experiments showed that the proposed method significantly improved the consistency between graphemes and the predicted labels. Further, experiments on accent estimation task confirmed that the created parallel data by the proposed method effectively improve the estimation accuracy.

语音标注拼音对齐语音合成口音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。