arXiv:2510.16813eess.AS2025-10

用相位感知正则化修复音频量化失真,避免能量损失。

Audio dequantization using instantaneous frequency

  • 引入相位感知正则化,保持时频域正弦分量的时序连续性
  • 在SDR和PEMO-Q指标上优于当前最优方法,主观听感更自然
  • 适合需要高保真音频重建的场景,如音乐修复与编解码

本文提出一种基于相位感知正则化的音频去量化方法(PHADQ),该正则化项最初成功应用于音频补全任务。该方法通过增强音频信号在时频表示中正弦分量的时序连续性,有效避免了传统L1正则化常导致的能量损失伪影。我们在SDR和PEMO-Q ODG客观指标下与现有最优方法对比,并进行了类MUSHRA的主观测试。结果表明,PHADQ在客观指标和主观听感上均表现更优,尤其在低比特率音频重建中展现出更强的鲁棒性。

原文摘要 · Abstract (English)

We present a dequantization method that employs a phase-aware regularizer, originally successfully applied in an audio inpainting problem. The method promotes a temporal continuity of sinusoidal components in time-frequency representation of the audio signal, and avoids energy loss artifacts commonly encountered with l1-based regularization approaches. The proposed method is called the Phase-Aware Audio Dequantizer (PHADQ). The method is evaluated against the state-of-the-art using the SDR and PEMO-Q ODG objective metrics, and a~subjective MUSHRA-like test.

音频重建去量化时频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。