arXiv:2608.19843cs.SDcs.MM2026-08

针对音乐重建中的高频丢失问题,提出频域感知自编码器提升音质。

Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction

论文配图:Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
图 1 · 摘自论文原文
  • 在频域显式建模,通过复数STFT实现逐频段精准控制
  • 在546首曲目数据集上,多项指标优于现有方法19.4%以上
  • 适合音频生成、音乐修复等需高保真还原的场景

连续潜空间音频自编码器是潜在音乐生成器的核心,但在高压缩率下常出现高频损失、相位不一致和立体声图像坍塌三种失效模式。其根源在于波形自编码器缺乏显式的频率轴,无法对各频段进行针对性修正。在五种预算相当的表示中,复数STFT在全频带和高频谱距离上表现最佳,可直接访问每个频点的幅度与相位。基于此,我们提出ear-VAE2,一种具有跨通道交互的复频谱自编码器。Spec-SnakeBeta在每个频率点学习周期性激活,使用更少参数且优于其他激活方式。Duplex-Aware Refiner根据双耳听觉理论对幅度和相位进行分频段修正。在546首曲目的Song Describer Dataset上,ear-VAE2在七项重建指标中有五项取得最优结果。该重构器使梅尔距离降低19.4%,残差输出维度减少约45%,同时降低谱距离、空间线索误差,并获得专业工程师更高评分。下游生成器使用ear-VAE2潜变量后,在全部12项自动评估指标上均表现更优。

原文摘要 · Abstract (English)

Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget representations, the complex STFT achieves the lowest full-band and high-frequency spectral distances, providing direct access to magnitude and phase at every bin. Building on this, we present ear-VAE2, a complex-spectral autoencoder with cross-channel interaction. Spec-SnakeBeta learns a periodic activation per frequency bin with frequency-dependent initialization, outperforming other activation variants while using fewer parameters than the fully independent variant. Duplex-Aware Refiner applies band-specific corrections to magnitude and phase following duplex theory of sound localization. On the 546-track Song Describer Dataset, ear-VAE2 achieves the best point estimates on five of seven reconstruction metrics. The Duplex-Aware Refiner reduces Mel Distance by 19.4% and uses ~45% fewer residual-output dimensions than the Unconstrained Refiner, while also lowering spectral distances, spatial-cue errors, and receiving higher ratings from professional engineers. The downstream generator using ear-VAE2 latents achieves better point estimates on all 12 automatic metrics.Demo page is available at https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/.

音频生成频域建模自编码器高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。