arXiv:2606.05522cs.SDcs.AI2026-06

首次系统评估大模型对南亚古典音乐的理解与生成能力,发现其表现受限于文化结构差异。

Exploring LLMs for South Asian Music Understanding and Generation

论文配图:Exploring LLMs for South Asian Music Understanding and Generation
图 1 · 摘自论文原文
  • 基于拉格与塔拉体系构建评估基准,测试大模型对南亚音乐规则的理解
  • 顶尖模型在问答任务中准确率达85-90%,但生成仅40%符合风格要求
  • 揭示了结构性正确与风格忠实是两个独立挑战,适合跨文化音乐研究者参考

近年来大型语言模型(LLMs)在音乐理解与生成任务中取得显著进展。然而现有工作仍局限于西方调性传统,对结构性迥异的低资源音乐传统缺乏探索。本文首次系统评估大模型在南亚古典音乐中的能力,该传统以拉格(raga)和塔拉(tala)为基础,遵循与西方和声驱动音乐截然不同的结构原则。评估基于印度斯坦古典理论及孟加拉古典形式(包括罗宾德拉与纳兹鲁尔·桑吉特),代表南亚古典音乐中的低资源传统。针对音乐理解,我们构建了一个包含504个问题的答案基准,涵盖拉格语法、文化知识与符号记谱推理,评估了33个大模型,其中前沿模型如Gemini 2.5 Pro达到85-90%准确率,而多数开源模型仅在23-40%之间。在音乐生成方面,我们设计五级受控提示框架,发现即使最强模型也仅在40%情况下生成风格一致的结果。结果表明,结构有效性与风格忠实是可分离的目标,凸显文化根基型音乐建模的开放挑战。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have shown promising results in music understanding and generation tasks. However, existing works remain confined to Western tonal traditions, offering little insight into whether current LLMs can handle structurally distinct low-resource musical traditions. We present the first systematic evaluation of LLM competence in South Asian classical music, a tradition governed by raga, tala-based melodic constraints that impose fundamentally different structural principles from Western harmony-driven music. We ground our evaluation in Hindustani classical theory and Bengali classical forms, including Rabindra and Nazrul Sangeet -- representative low-resource traditions within South Asian classical music. For music understanding evaluation, we introduce a 504-question-answer benchmark spanning raga grammar, cultural knowledge, and symbolic notation reasoning, evaluating 33 LLMs where frontier models such as Gemini 2.5 Pro achieve 85-90% accuracy, while most open-source models remain in the 23-40% range. For music generation, we design a five-level controlled prompting framework and find that even the strongest model produces stylistically faithful outputs only 40% of the time. These results reveal that structural validity and stylistic faithfulness in music generation are distinct objectives and highlight an open challenge for culturally grounded music modeling.

音乐理解大模型南亚音乐文化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。