构建跨文化低资源音乐理解基准,提升模型对民间音乐的感知能力
UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

- 设计自动化流程构建5042组跨文化音乐问答数据
- 用自动生成对话数据训练模型,显著改善多模态不平衡问题
- 适合关注文化多样性与音频理解的AI研究者
大型音频-语言模型在音乐描述、流派分类和声音事件检测任务中表现优异,但对根植于不同文化背景的民间音乐适应性不足。这类音乐普遍资源稀缺、分布不均且记录不全,即使出现在大规模预训练中,模型也难以捕捉其结构与风格特征,部分原因在于缺乏专门的评估协议与训练方案。为此,我们提出UniVerse,一套可复现的低资源音乐理解解决方案。具体包括:基于专家指导的自动化流程构建的UniVerseBench基准,包含5042个跨38个文化语言实体的问答对;以及完全自动化的多轮对话训练数据集UniVerseSet。通过在UniVerseSet上训练,系统评估了多种主流多模态不平衡学习策略在密集型与混合专家(MoE)架构中的效果。实验表明,全自动数据构建结合失衡感知训练带来显著提升,但模型仍难以捕捉精细声学特征,说明表面对齐与深层音乐理解之间存在明显差距。
原文摘要 · Abstract (English)
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。