arXiv:2605.03929cs.SDcs.AI2026-05中稿 · ICML被引 1

用相位信息提升音乐音频分离准确率,更省资源更快

PHALAR: Phasors for Learned Musical Audio Representations

论文配图:PHALAR: Phasors for Learned Musical Audio Representations
图 1 · 摘自论文原文
  • 引入复数域表示与谱池化层,保留音高和相位等变性
  • 在三个数据集上准确率提升约70%,参数量减半,训练快7倍
  • 不仅适合音频分离,还能零样本做节拍追踪和和弦分析

声部检索(stem retrieval)是将缺失的声部与给定音频子混合匹配的关键挑战,当前模型常因丢失时间信息而受限。我们提出PHALAR,一种对比学习框架,在保持低于50%参数量的同时,实现约70%的相对准确率提升,并带来7倍的训练加速。通过采用学习型谱池化层和复数输出头,PHALAR 强制引入音高等变性和相位等变性先验。该方法在MoisesDB、Slakh和ChocoChorales上均建立新的检索性能标杆,且与人类对连贯性的判断相关性显著高于语义基线。此外,零样本节拍追踪和线性和弦探测结果表明,PHALAR还捕捉到了超越检索任务的稳健音乐结构。

原文摘要 · Abstract (English)

Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a contrastive framework achieving a relative accuracy increase of up to $\approx 70\%$ over the state-of-the-art while requiring $<50\%$ of the parameters and a 7$\times$ training speedup. By utilizing a Learned Spectral Pooling layer and a complex-valued head, PHALAR enforces pitch-equivariant and phase-equivariant biases. PHALAR establishes new retrieval state-of-the-art across MoisesDB, Slakh, and ChocoChorales, correlating significantly higher with human coherence judgment than semantic baselines. Finally, zero-shot beat tracking and linear chord probing confirm that PHALAR captures robust musical structures beyond the retrieval task.

音乐生成音频分离复数表示等变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。