用流模型将干净语音转为高斯分布,精准引导语音修复
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
- 通过流模型将干净语音映射到预设高斯分布,提取引导信息
- 在生成网络每层嵌入引导信号,实现精细修复控制
- 适用于各类通用语音修复任务,提升恢复质量
我们提出FLOWER,一种新型语音修复条件方法,将高斯引导融入生成框架。通过归一化流网络将干净语音转换为预设先验分布(如高斯分布),提取关键信息以指导生成模型。该引导信息被嵌入生成网络的每个模块,实现精确的修复控制。实验表明,FLOWER在多种通用语音修复任务中均有效提升了性能。
原文摘要 · Abstract (English)
We introduce FLOWER, a novel conditioning method designed for speech restoration that integrates Gaussian guidance into generative frameworks. By transforming clean speech into a predefined prior distribution (e.g., Gaussian distribution) using a normalizing flow network, FLOWER extracts critical information to guide generative models. This guidance is incorporated into each block of the generative network, enabling precise restoration control. Experimental results demonstrate the effectiveness of FLOWER in improving performance across various general speech restoration tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。