用音乐特征辅助生成更精准的音乐描述,提升音轨理解能力。
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
- 通过多任务学习融合歌词生成与关键音乐特征检测
- 结合声学细节与高阶音乐属性,生成更丰富的描述
- 适合音乐信息检索、智能推荐等应用
准确反映音乐作品特征的详细描述能丰富音乐数据库并推动音乐AI研究。本文提出多任务音乐描述模型SonicVerse,将文本生成与关键检测、人声识别等辅助音乐特征任务联合训练,直接捕捉低层声学细节和高层音乐属性。核心是基于投影的架构:音频输入被转换为语言令牌,同时通过专用辅助头检测音乐特征,这些特征输出也被投影为语言令牌以增强描述输入。该框架不仅能为短音乐片段生成丰富描述,还可通过大语言模型串联输出,实现对长音乐段落的时间感知描述。为训练模型,我们利用MIRFLEX模块化音乐特征提取器扩展了MusicBench数据集,标注了配对的音频、描述与音乐特征数据。实验表明,引入特征显著提升了生成描述的质量与细节。
原文摘要 · Abstract (English)
Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, that integrates caption generation with auxiliary music feature detection tasks such as key detection, vocals detection, and more, so as to directly capture both low-level acoustic details as well as high-level musical attributes. The key contribution is a projection-based architecture that transforms audio input into language tokens, while simultaneously detecting music features through dedicated auxiliary heads. The outputs of these heads are also projected into language tokens, to enhance the captioning input. This framework not only produces rich, descriptive captions for short music fragments but also directly enables the generation of detailed time-informed descriptions for longer music pieces, by chaining the outputs using a large-language model. To train the model, we extended the MusicBench dataset by annotating it with music features using MIRFLEX, a modular music feature extractor, resulting in paired audio, captions and music feature data. Experimental results show that incorporating features in this way improves the quality and detail of the generated captions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。