用相位先验修复音频频谱缺失,效果好且速度快
Audio Inpainting in Time-Frequency Domain with Phase-Aware Prior
- 基于瞬时频率估计构建相位感知先验
- 短缺口下信噪比优于现有方法,长缺口持平最优
- 计算开销远低于深度模型,适合实时应用
针对时频域音频补全问题,本文提出一种利用瞬时频率估计构建相位感知信号先验的方法。通过构造优化问题并采用广义Chambolle-Pock算法求解,实现对缺失频谱片段的一致性重建。在短间隙情况下,该方法在客观指标上优于现有方法;在长间隙下性能与基于自回归的Janssen-TF方法相当。主观与客观听觉评估均显示,该方法在所有间隙长度下表现更优。此外,相比其他方法,其计算成本显著降低。
原文摘要 · Abstract (English)
We address the problem of time-frequency audio inpainting, where the goal is to fill missing spectrogram portions with consistent information. Despite recent advances, existing approaches still face limitations in both reconstruction quality and computational efficiency. To bridge this gap, we propose a method that utilizes a phase-aware signal prior which exploits estimates of the instantaneous frequency. An optimization problem is formulated and solved using the generalized Chambolle-Pock algorithm. The proposed method is evaluated against other time-frequency inpainting methods, specifically a deep-prior audio inpainting neural network, the autoregression-based approach known as Janssen-TF, and a sparsity-driven baseline. For short gap durations, the proposed approach achieves superior SNR, while performing comparably to Janssen-TF on larger gaps. In terms of perceptual quality (both objective and subjective), the proposed method consistently outperforms existing methods across all gap lengths. In addition, the reconstructions are obtained with a substantially reduced computational cost compared to alternative methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。