arXiv:2509.14912cs.SDcs.AI2025-09被引 7

针对音乐重建中的听觉感知缺陷,提出新型VAE模型提升音高精度与立体声效果。

Back to Ear: Perceptually Driven High Fidelity Music Reconstruction

  • 引入感知加权滤波器,使损失函数更贴近人耳听觉特性。
  • 设计相位相关性损失与瞬时频率损失,显著提升相位精度。
  • 仅用左右通道监督相位,提升高频谐波和空间感重建质量。

变分自编码器(VAEs)在大规模音频任务中至关重要,但现有开源模型常忽视听觉感知,导致相位准确性和立体声空间表征弱。为此,我们提出εar-VAE,一种重新设计训练范式的开源音乐信号重建模型。贡献包括:(i) 在损失计算前应用K-weighting感知滤波器,使目标对齐听觉感知;(ii) 提出两种新相位损失:用于立体声一致性的相关性损失,以及基于瞬时频率和群延迟的相位损失;(iii) 新谱监督范式:幅度由中/侧/左/右四通道共同监督,相位仅由左右通道监督。实验表明,44.1kHz下的εar-VAE在多项指标上显著优于主流开源模型,尤其在高频频谱谐波和空间特性重建方面表现突出。

原文摘要 · Abstract (English)

Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual aspects during training, leading to weaknesses in phase accuracy and stereophonic spatial representation. To address these challenges, we propose εar-VAE, an open-source music signal reconstruction model that rethinks and optimizes the VAE training paradigm. Our contributions are threefold: (i) A K-weighting perceptual filter applied prior to loss calculation to align the objective with auditory perception. (ii) Two novel phase losses: a Correlation Loss for stereo coherence, and a Phase Loss using its derivatives--Instantaneous Frequency and Group Delay--for precision. (iii) A new spectral supervision paradigm where magnitude is supervised by all four Mid/Side/Left/Right components, while phase is supervised only by the LR components. Experiments show εar-VAE at 44.1kHz substantially outperforms leading open-source models across diverse metrics, showing particular strength in reconstructing high-frequency harmonics and the spatial characteristics.

音乐重建感知建模相位优化VAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。