arXiv:2603.08046cs.SDeess.AS2026-03

用双向框架生成伪平行数据,提升耳语转正常语音质量

WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation

  • 基于统一语义表征捕捉耳语与正常语音共性特征
  • 通过零样本生成伪耳语数据,显著提升转换性能
  • 适合语音转换、数据增强研究者参考

耳语因缺乏声带振动和基频,导致声学线索退化,使得耳语转正常语音(W2N)在并行数据有限的情况下尤为困难。本文提出 WhispEar,一种基于统一语义表征的双向框架,可捕获耳语与正常语音共享的说话模式不变信息。该框架包含 W2N 和正常语音转耳语(N2W)模型。其中 N2W 模型可从大量正常语音中实现零样本伪耳语生成,从而实现可扩展的数据增强,用于训练 W2N 模型。生成数据越多,性能越优。同时,本文发布了目前最大的双语(中英)耳语-正常语音并行语料库。实验表明,WhispEar 显著优于强基线方法,且对可扩展伪并行数据有显著收益。

原文摘要 · Abstract (English)

Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a bidirectional framework based on unified semantic representations that capture speaking-mode-invariant information shared by whispered and normal speech. The framework contains both W2N and normal-to-whisper (N2W) models. Notably, the N2W model enables zero-shot pseudo-parallel whisper generation from abundant normal speech, allowing scalable data augmentation for W2N training. Increasing generated data consistently improves performance. We also release the largest bilingual (Chinese-English) whispered-normal parallel corpus to date. Experiments demonstrate that WhispEar outperforms strong baselines and benefits significantly from scalable pseudo-parallel data.

语音转换数据增强双向模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。