无需配对数据,用隐空间优化修复未知音频退化问题
Music Restoration via Latent Operator Optimization and Diffusion Model Priors
- 在预训练音频自编码器的隐空间中建模未知失真为可学习算子
- 通过交替优化隐变量与算子参数,实现对退化音频的恢复
- 结合无条件扩散模型先验,适用于多种未见退化场景
音乐修复旨在从受未知效应、失真或损坏影响的观测录音中恢复出干净信号。现有系统通常依赖成对训练数据和特定失真监督,当正向过程未知时难以应用。我们提出 LOUDAR(Latent-space Optimization of Unknown Distortion for Audio Restoration),一种通用修复方法,其在预训练音频自编码器的隐空间中运行,并将未知失真建模为可学习的隐空间算子。推理时,LOUDAR 交替估计干净隐变量并更新隐算子参数。一个无条件隐空间扩散模型提供干净音频先验,通过引导隐变量估计趋向干净录音流形来正则化推断过程。由于退化模型针对每输入自适应调整,该方法广泛适用于各类修复任务。我们在人声效果去除、歌唱语音修复及吉他失真去除上评估了 LOUDAR,结果表明其持续改善退化输入,在波形与隐空间上均达到与有监督和无监督基线相当的性能。
原文摘要 · Abstract (English)
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We propose LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) a general-purpose restoration method that operates in the latent space of a pretrained audio autoencoder and models the unknown distortion as a learnable latent operator. At inference time, LOUDAR alternates between estimating the clean latent variable and updating the latent operator parameters. An unconditional latent diffusion model provides a prior over clean audio and regularizes this inference by steering the latent estimate toward the manifold of clean recordings. Because the degradation model is adapted per input, the approach is broadly applicable across diverse restoration problems. We evaluate LOUDAR on singing voice effect removal and restoration, as well as guitar distortion removal, and show that it consistently improves over degraded inputs and is competitive with supervised and unsupervised baselines in waveform and latent domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。