arXiv:2502.06710cs.CVcs.MM2025-02EMNLP被引 29

为音乐表演问答设计更懂乐器和节奏的多模态模型

Learning Musical Representations for Music Performance Question Answering

  • 构建音乐场景下融合音视频交互的模型架构
  • 在现有数据集上实现音乐问答任务新最优性能
  • 适合研究音乐理解与多模态交互的学者

音乐表演是音频-视觉建模的典型场景,其持续密集的音频信号区别于普通稀疏音频场景。现有音视频问答方法虽在通用场景表现优异,但在音乐表演中因忽视多模态信号间的深层交互及乐器与音乐特性,导致回答不准确。为此,我们提出:(i) 设计以音乐上下文为基础的多模态交互主干网络;(ii) 在现有音乐数据集中标注并发布节奏与声源信息;(iii) 引入时间感知对齐机制,使模型预测与时间维度同步。实验表明,在Music AVQA数据集上达到当前最优效果。代码已开源:https://github.com/xid32/Amuse。

原文摘要 · Abstract (English)

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the audio-video QA demonstrate impressive capabilities in general scenarios, they are incapable of dealing with fundamental problems within the music performances: they underexplore the interaction between the multimodal signals in performance and fail to consider the distinctive characteristics of instruments and music. Therefore, existing methods tend to answer questions regarding musical performances inaccurately. To bridge the above research gaps, (i) given the intricate multimodal interconnectivity inherent to music data, our primary backbone is designed to incorporate multimodal interactions within the context of music; (ii) to enable the model to learn music characteristics, we annotate and release rhythmic and music sources in the current music datasets; (iii) for time-aware audio-visual modeling, we align the model's music predictions with the temporal dimension. Our experiments show state-of-the-art effects on the Music AVQA datasets. Our code is available at https://github.com/xid32/Amuse.

音乐理解多模态问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。