arXiv:2511.17555eess.AScs.CL2025-11AAAI被引 1

用语音识别模型的注意力机制,让语音合成更精准地对齐文字。

Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward

  • 利用预训练语音识别模型的注意力,实现逐词级反馈信号。
  • 在未见过的说话人上,语音合成质量与鲁棒性显著提升。
  • 无需人工标注奖励,适合改进现有语音合成系统。

近年来,文本到语音(TTS)技术已能克隆任意未见说话人并生成高质量、自然的语音。然而,评估方法滞后:传统均值意见分数(MOS)估算器对整个语句进行回归,而错误通常只出现在少数词语上。我们发现编码器-解码器型语音识别模型(如Whisper)通过交叉注意力可揭示语音与文本之间的词级不匹配,从而提供细粒度奖励信号。基于此,我们提出一种由语音识别驱动的注意力奖励机制(W3AR),无需显式奖励标注,即可利用预训练语音识别模型的注意力,引导TTS模型预测序列的细粒度对齐与优化。实验表明,W3AR能提升现有TTS系统的质量,并增强对未见说话人的零样本鲁棒性。更广泛而言,结果提示一种生成建模的简单范式:理解模型可充当评估器,为优化提供信息丰富、细粒度的反馈。

原文摘要 · Abstract (English)

Recent advances in text-to-speech (TTS) have enabled models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, evaluation methods lag behind: typical mean opinion score (MOS) estimators perform regression over entire utterances, while failures usually occur in a few problematic words. We observe that encoder-decoder ASR models (e.g., Whisper) surface word-level mismatches between speech and text via cross-attention, providing a fine-grained reward signal. Building on this, we introduce Word-level TTS Alignment by ASR-driven Attentive Reward (W3AR). Without explicit reward annotations, W3AR uses attention from a pre-trained ASR model to drive finer-grained alignment and optimization of sequences predicted by a TTS model. Experiments show that W3AR improves the quality of existing TTS systems and strengthens zero-shot robustness on unseen speakers. More broadly, our results suggest a simple recipe for generative modeling: understanding models can act as evaluators, delivering informative, fine-grained feedback for optimization.

语音合成注意力机制评估方法零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。