用合成数据训练,让耳语变正常语音,识别更准
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
- 用流匹配方法,基于合成的对齐耳语-语音对训练
- 在真实耳语上实现顶尖识别率,错词率降低26%-46%
- 无需真实配对数据,适合语音重建与无障碍应用
耳语转正常语音转换旨在从耳语输入中重建缺失的声带振动信息,同时保持内容和说话人身份。该任务因耳语与有声录音在时间上不对齐且缺乏成对数据而具有挑战性。我们提出 FlowW2N,一种仅在合成、时间对齐的耳语-正常语音对上训练的条件流匹配方法,并利用域不变特征进行条件控制。我们利用高阶ASR嵌入,其在合成与真实耳语间表现出强不变性,从而实现对真实耳语的泛化,尽管训练时从未见过真实耳语。我们在ASR各层验证了这种不变性,并提出了优化内容信息量与跨域不变性的选择准则。所提方法在CHAINS和wTIMIT数据集上达到最佳可懂度,相较于之前工作,相对错词率降低26%-46%,推理仅需10步,且无需真实配对数据。
原文摘要 · Abstract (English)
Whispered-to-normal (W2N) speech conversion aims to reconstruct missing phonation from whispered input while preserving content and speaker identity. This task is challenging due to temporal misalignment between whisper and voiced recordings and lack of paired data. We propose FlowW2N, a conditional flow matching approach that trains exclusively on synthetic, time-aligned whisper-normal pairs and conditions on domain-invariant features. We exploit high-level ASR embeddings that exhibits strong invariance between synthetic and real whispered speech, enabling generalization to real whispers despite never observing it during training. We verify this invariance across ASR layers and propose a selection criterion optimizing content informativeness and cross-domain invariance. Our method achieves SOTA intelligibility on the CHAINS and wTIMIT datasets, reducing Word Error Rate by 26-46% relative to prior work while using only 10 steps at inference and requiring no real paired data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。