GaMMA统一时序与非时序音乐理解,实现音频-文本跨模态新标杆
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

- 采用专家混合架构融合时序与非时序音乐特征,共享同一参数集
- 在三个任务上达到79.1%~81.3%准确率,超越现有方法
- 适用于音乐理解、跨模态生成等需要全局与时间感知的场景
本文提出GaMMA,一种先进大型多模态模型(LMM),旨在实现全面的音乐内容理解。该模型沿用LLaVA的轻量编码器-解码器结构,有效实现音乐与语言间的跨模态学习。通过以专家混合方式引入音频编码器,GaMMA在单一参数集下统一处理时序与非时序音乐理解任务。研究结合大规模精心筛选数据集与渐进式训练流程,通过预训练、监督微调(SFT)和强化学习(RL)持续提升音乐理解能力。为全面评估音乐类LMM的时间与非时间能力,我们构建了目前最大的音乐导向基准MusicBench,包含3,739道人工精选的多项选择题,覆盖多样音乐理解维度。大量实验表明,GaMMA在音乐领域建立新基准,于MuchoMusic上达79.1%准确率,MusicBench-Temporal达79.3%,MusicBench-Global达81.3%,持续优于先前方法。
原文摘要 · Abstract (English)
In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, enabling effective cross-modal learning between music and language. By incorporating audio encoders in a mixture-of-experts manner, GaMMA effectively unifies both time-series and non-time-series music understanding tasks within one set of parameters. Our approach combines carefully curated datasets at scale with a progressive training pipeline, effectively pushing the boundaries of music understanding via pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL). To comprehensively assess both temporal and non-temporal capability of music LMMs, we introduce MusicBench, the largest music-oriented benchmark, comprising 3,739 human-curated multiple-choice questions covering diverse aspects of musical understanding. Extensive experiments demonstrate that GaMMA establishes new SoTA in the music domain, achieving 79.1% accuracy on MuchoMusic, 79.3% on MusicBench-Temporal, and 81.3% on MusicBench-Global, consistently outperforming previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。