arXiv:2409.08795eess.AScs.MM2024-09被引 11

用大模型分析音乐表演,给出个性化反馈。

LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment

  • 基于音频语言建模,通过提问方式评估演奏表现
  • 在教师评分预测、曲目难度识别上达顶尖水平
  • 适合音乐教育、智能伴奏系统开发者使用

音乐理解研究已通过高级表征广泛探索了调性、风格和配器等作曲层面属性,并推动了跨模态应用的发展。然而,演奏中的风格表达与技巧等维度仍研究不足,且大语言模型在提供定制化反馈以提升教学效果方面的潜力尚未被充分挖掘。为此,我们提出LLaQo——一种基于大语言模型的音乐演奏问答教练,利用音频语言建模对音乐表演进行详细而具有形成性的评估。我们还构建了指令微调的问答数据集,涵盖从音高准确性到演奏技法在内的多种表演维度,以及情境化理解(如曲目难度和演奏技巧)。采用AudioMAE编码器与Vicuna-7b LLM后端,该模型在预测教师评分、识别曲目难度与演奏技法方面均达到当前最优(SOTA)表现。用户研究显示,相较于其他基线模型,LLaQo生成的文本回应在音频-文本匹配任务中获得显著更高评价。因此,该模型可基于音频数据对音乐演奏相关的开放问题提供信息丰富、准确的回答。

原文摘要 · Abstract (English)

Research in music understanding has extensively explored composition-level attributes such as key, genre, and instrumentation through advanced representations, leading to cross-modal applications using large language models. However, aspects of musical performance such as stylistic expression and technique remain underexplored, along with the potential of using large language models to enhance educational outcomes with customized feedback. To bridge this gap, we introduce LLaQo, a Large Language Query-based music coach that leverages audio language modeling to provide detailed and formative assessments of music performances. We also introduce instruction-tuned query-response datasets that cover a variety of performance dimensions from pitch accuracy to articulation, as well as contextual performance understanding (such as difficulty and performance techniques). Utilizing AudioMAE encoder and Vicuna-7b LLM backend, our model achieved state-of-the-art (SOTA) results in predicting teachers' performance ratings, as well as in identifying piece difficulty and playing techniques. Textual responses from LLaQo was moreover rated significantly higher compared to other baseline models in a user study using audio-text matching. Our proposed model can thus provide informative answers to open-ended questions related to musical performance from audio data.

音乐生成大模型教育应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。