将自回归音频修复方法拓展至时频域,效果优于现有深度先验模型。
Janssen 2.0: Audio Inpainting in the Time-frequency Domain
- 将时域自回归算法 adapted 到时频域,直接补全谱图缺失系数。
- 在客观指标与主观听感测试中均显著优于深度先验模型。
- 适合需要高质量音频修复的音乐处理、语音恢复场景。
本文聚焦于音频信号谱图中缺失部分的修复,即估计缺失的时频域系数。将基于自回归的先进时域音频修复方法 Janssen 算法适配至时频域,提出新方法 Janssen-TF。通过客观指标和主观听觉测试,与基于深度先验的神经网络方法进行对比,结果表明 Janssen-TF 在所有评估维度上均表现更优。
原文摘要 · Abstract (English)
The paper focuses on inpainting missing parts of an audio signal spectrogram, i.e., estimating the lacking time-frequency coefficients. The autoregression-based Janssen algorithm, a state-of-the-art for the time-domain audio inpainting, is adapted for the time-frequency setting. This novel method, termed Janssen-TF, is compared with the deep-prior neural network approach using both objective metrics and a subjective listening test, proving Janssen-TF to be superior in all the considered measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。