arXiv:2511.01056eess.AS2025-11中稿 · Interspeech 2026

将低资源耳语转为正常语音,分离对齐与生成步骤提升可懂度。

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

  • 分三阶段:先学语义不变表示,再建音色韵律模型,最后合成波形。
  • 在AISHELL6-Whisper数据集上实现CER 16.93%、DNSMOS 3.07、WavLM相似度0.95。
  • 适合隐私通信、无声交流及声带术后患者康复使用。

耳语因缺乏声带激励,难以实现可懂转换。我们提出WhisperVC,一种三阶段框架,用于低资源耳语转正常语音(W2N)转换,将跨域对齐与语音生成解耦。第一阶段利用少量成对耳语-正常语音数据,结合内容编码器与基于Conformer的变分自编码器(VAE),通过soft-DTW对齐学习域不变语义表征。第二阶段仅在正常语音上训练,采用长度-通道对齐器和两阶段说话人条件化的梅尔频谱生成器,建模音色与语调。第三阶段微调HiFi-GAN声码器完成波形合成。在AISHELL6-Whisper数据集上的实验表明,该方法达到竞争力的音质(DNSMOS 3.07,UTMOS 2.83,CER 16.93%)和说话人相似性(WavLM 0.95)。该框架还可用于隐私保护通信、非发声交流以及声带术后患者康复。样本已公开。

原文摘要 · Abstract (English)

Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) conversion that decouples cross-domain alignment from speech generation. Stage 1 uses limited paired whisper-normal data with a content encoder and a Conformer-based variational autoencoder (VAE) with soft-DTW alignment to learn domain-invariant semantic representations. Stage 2, trained only on normal speech, employs a Length-Channel Aligner and a two-stage speaker-conditioned mel generator for timbre and prosody modeling. Stage 3 fine-tunes a HiFi-GAN vocoder for waveform synthesis. Experimental results on AISHELL6-Whisper show competitive quality (DNSMOS 3.07, UTMOS 2.83, CER 16.93%) and WavLM speaker similarity (0.95). The framework also supports privacy-preserving communication as well as non-vocal communication and a rehabilitation tool for post-surgical vocal-fold patients. Samples are available online.

语音转换低资源耳语合成声带康复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。