arXiv:2601.14472cs.SDcs.AI2026-01中稿 · presentation at IC…

通过声调引导的谐波注意力提升语音合成的音高准确性和相位一致性。

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

  • 用声调引导的谐波注意力增强有声段表征,直接预测复频谱重建波形。
  • 在基准数据集上音高均方误差降22%,有声/无声识别错误降18%。
  • 适合关注音色自然度和音高保真的语音合成研究者与开发者。

神经声码器是语音合成的核心;尽管取得成功,多数仍存在声调建模有限和相位重建不准确的问题。本文提出一种新声码器,引入声调引导的谐波注意力以增强有声段编码,并通过逆短时傅里叶变换直接预测复频谱成分进行波形合成。不同于基于梅尔频谱的方法,该设计联合建模幅度与相位,确保相位一致性并提升音高保真度。为进一步匹配感知质量,采用多目标训练策略,融合对抗损失、谱损失和相位感知损失。在基准数据集上的实验表明,相比HiFi-GAN和AutoVocoder,F0 RMSE降低22%,有声/无声错误率下降18%,MOS评分提升0.15。结果表明,声调引导注意力结合直接复频谱建模可生成更自然、音高准确且鲁棒的合成语音,为表达性神经声码器奠定坚实基础。

原文摘要 · Abstract (English)

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance voiced segment encoding and directly predicts complex spectral components for waveform synthesis via inverse STFT. Unlike mel-spectrogram-based approaches, our design jointly models magnitude and phase, ensuring phase coherence and improved pitch fidelity. To further align with perceptual quality, we adopt a multi-objective training strategy that integrates adversarial, spectral, and phase-aware losses. Experiments on benchmark datasets demonstrate consistent gains over HiFi-GAN and AutoVocoder: F0 RMSE reduced by 22 percent, voiced/unvoiced error lowered by 18 percent, and MOS scores improved by 0.15. These results show that prosody-guided attention combined with direct complex spectrum modeling yields more natural, pitch-accurate, and robust synthetic speech, setting a strong foundation for expressive neural vocoding.

语音合成声码器复频谱音高建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。