用双向框架生成伪平行数据,提升耳语转正常语音质量
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
- 基于统一语义表征捕捉耳语与正常语音共性特征
- 通过零样本生成伪耳语数据,显著提升转换性能
- 适合语音转换、数据增强研究者参考
耳语因缺乏声带振动和基频,导致声学线索退化,使得耳语转正常语音(W2N)在并行数据有限的情况下尤为困难。本文提出 WhispEar,一种基于统一语义表征的双向框架,可捕获耳语与正常语音共享的说话模式不变信息。该框架包含 W2N 和正常语音转耳语(N2W)模型。其中 N2W 模型可从大量正常语音中实现零样本伪耳语生成,从而实现可扩展的数据增强,用于训练 W2N 模型。生成数据越多,性能越优。同时,本文发布了目前最大的双语(中英)耳语-正常语音并行语料库。实验表明,WhispEar 显著优于强基线方法,且对可扩展伪并行数据有显著收益。
原文摘要 · Abstract (English)
Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a bidirectional framework based on unified semantic representations that capture speaking-mode-invariant information shared by whispered and normal speech. The framework contains both W2N and normal-to-whisper (N2W) models. Notably, the N2W model enables zero-shot pseudo-parallel whisper generation from abundant normal speech, allowing scalable data augmentation for W2N training. Increasing generated data consistently improves performance. We also release the largest bilingual (Chinese-English) whispered-normal parallel corpus to date. Experiments demonstrate that WhispEar outperforms strong baselines and benefits significantly from scalable pseudo-parallel data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。