用流匹配实现单步可控音乐频宽扩展,提升修复精度。
Single-step Controllable Music Bandwidth Extension With Flow Matching
- 基于流匹配框架,引入动态频谱轮廓控制信号
- 在频宽扩展任务中表现优于现有方法,支持精细调控
- 适合需要高保真音频修复的音乐档案场景
音频恢复旨在逆向数字音频信号的退化过程,还原其原始清晰版本,对具有珍贵历史价值的音乐录音档案尤为重要。近年来,生成模型在音频恢复中展现出显著优势,突破了传统方法的局限。然而,如何实现模型的精细可控仍是难题。本文扩展了FLowHigh模型,提出动态频谱轮廓(DSC)作为无分类器引导下的频宽扩展控制信号。实验表明,该方法在性能上具有竞争力,且DSC能有效支持细粒度条件控制,为高质量音频重建提供新路径。
原文摘要 · Abstract (English)
Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is valuable in contexts such as archives of music recordings, particularly those of precious historical value, for which a clean version may have been lost or simply does not exist. Recent work applied generative models to audio restoration, showing promising improvement over previous methods, and opening the door to the ability to perform restoration operations that were not possible before. However, making these models finely controllable remains a challenge. In this paper, we propose an extension of FLowHigh and introduce the Dynamic Spectral Contour (DSC) as a control signal for bandwidth extension via classifier-free guidance. Our experiments show competitive model performance, and indicate that DSC is a promising feature to support fine-grained conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。