用声音环境信息增强语音修复,让模型更懂噪声来源。
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
- 用CLAP提取音频环境特征,结合扩散模型进行修复
- 新提出的ACX表示能更好应对不同强度的失真
- 适合处理复杂噪声场景下的语音恢复任务
本文提出一种基于声学上下文表征的生成式语音修复方法。以基于扩散模型的UNIVERSE++为基线,引入来自CLAP模型的声学上下文嵌入,捕捉输入音频的环境属性。同时提出一种改进的声学上下文(ACX)表示,对CLAP嵌入进行优化,以更有效地应对各类失真及其强度变化。与依赖语言和说话人属性的内容型方法不同,ACX提供上下文信息,使修复模型能更准确地区分并抑制失真。实验表明,采用上下文感知条件化后,修复性能与稳定性均提升,且在多种失真条件下波动更小。
原文摘要 · Abstract (English)
This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the effectiveness of contextual representations. We incorporate acoustic context embeddings extracted from the CLAP model, which capture the environmental attributes of input audio. Additionally, we propose an Acoustic Context (ACX) representation that refines CLAP embeddings to better handle various distortion factors and their intensity in speech signals. Unlike content-based approaches that rely on linguistic and speaker attributes, ACX provides contextual information that enables the restoration model to distinguish and mitigate distortions better. Experimental results indicate that context-aware conditioning improves both restoration performance and its stability across diverse distortion conditions, reducing variability compared to content-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。