arXiv:2509.11976cs.SDeess.AS2025-09

用压缩音频特征提升音乐情感分析的多模态融合效果

PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis

  • 用向量量化+空间池化压缩音频特征,减少冗余
  • 在EMOPIA和VGMIDI数据集上达到最新最好性能
  • 适合做音乐情感分析或多模态融合研究者参考

多模态音乐情感分析通过结合音频与MIDI模态提升性能。主流方法聚焦于复杂特征提取网络,而我们提出:通过缩短音频序列特征长度来缓解冗余,尤其对比紧凑的MIDI表示,可有效提升任务表现。为此,我们设计了PoolingVQ,将向量量化变分自编码器(VQVAE)与空间池化结合,通过码本引导的局部聚合直接压缩音频特征序列以降低冗余;同时提出两阶段协同注意力机制融合音频与MIDI信息。在公开数据集EMOPIA和VGMIDI上的实验表明,该多模态框架达到当前最优性能,PoolingVQ显著提升效果。代码已开源于匿名GitHub。

原文摘要 · Abstract (English)

Multimodal music emotion analysis leverages both audio and MIDI modalities to enhance performance. While mainstream approaches focus on complex feature extraction networks, we propose that shortening the length of audio sequence features to mitigate redundancy, especially in contrast to MIDI's compact representation, may effectively boost task performance. To achieve this, we developed PoolingVQ by combining Vector Quantized Variational Autoencoder (VQVAE) with spatial pooling, which directly compresses audio feature sequences through codebook-guided local aggregation to reduce redundancy, then devised a two-stage co-attention approach to fuse audio and MIDI information. Experimental results on the public datasets EMOPIA and VGMIDI demonstrate that our multimodal framework achieves state-of-the-art performance, with PoolingVQ yielding effective improvement. Our proposed metho's code is available at Anonymous GitHub

音频处理多模态情感分析VQVAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。