arXiv:2603.00533cs.SDeess.AS2026-03中稿 · ISMIR 2025 LBD被引 1

首个多语言音乐理解问答基准,测试大模型对全球音乐文化的理解能力。

Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding

  • 构建覆盖38种语言的音乐问答数据集,自动化生成1190道多选题。
  • 顶尖音频大模型在无丰富文本背景时,难以捕捉细微文化差异。
  • 适合研究跨文化音频理解、公平性评估的研究者使用。

我们提出 Voices of Civilizations,首个用于评估音频大模型在完整音乐录音上文化理解能力的多语言问答基准。该数据集涵盖380首来自38种语言的曲目,通过四阶段自动化流程生成1,190道多选题:1)编制代表性音乐列表;2)利用大模型生成每首曲目的文化背景文档;3)从文档中提取关键属性;4)构造探查语言、地域关联、情绪与主题内容的问题。我们在四种条件下评估模型并报告各语言准确率。结果表明,即使最先进的音频大模型在缺乏丰富文本上下文时,也难以把握细微文化内涵,并在解读不同文化传统音乐时表现出系统性偏差。该数据集已公开发布于 Hugging Face,以推动更具包容性的音乐理解研究。

原文摘要 · Abstract (English)

We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list; 2) generating cultural-background documents for each sample in the music list via LLMs; 3) extracting key attributes from those documents; and 4) constructing multiple-choice questions probing language, region associations, mood, and thematic content. We evaluate models under four conditions and report per-language accuracy. Our findings demonstrate that even state-of-the-art audio LLMs struggle to capture subtle cultural nuances without rich textual context and exhibit systematic biases in interpreting music from different cultural traditions. The dataset is publicly available on Hugging Face to foster culturally inclusive music understanding research.

音乐理解多语言文化偏见音频LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。