arXiv:2410.03264cs.SDcs.IR2024-10中稿 · publication at the…被引 23

用微调大模型生成丰富音乐描述,提升文本搜歌准确率

Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval

  • 用微调LLM和元数据生成更丰富的音乐描述文本
  • 在多种查询下优于现有顶尖联合嵌入模型
  • 适合需要根据风格或情感搜歌的用户

文本到音乐检索在大规模音乐数据库中对内容发现至关重要。现有研究多聚焦于音乐音频与文本的联合嵌入,以匹配与音乐属性(如流派、乐器)和上下文元素(如情绪、主题)相关的精确描述。然而,用户也常希望找到与喜爱曲目或艺人相似的音乐,例如“找一首像斯蒂维·旺德《Superstition》的歌曲”。为此,本文提出改进的文本到音乐检索模型TTMR++,利用微调大语言模型生成的丰富文本描述及元数据。我们从多个音乐标签与标题数据集以及艺术家和曲目知识图谱中获取各类初始文本。实验表明,在涵盖多种音乐文本查询的综合评估中,TTMR++显著优于当前最先进的音乐-文本联合嵌入模型。

原文摘要 · Abstract (English)

Text-to-Music Retrieval, finding music based on a given natural language query, plays a pivotal role in content discovery within extensive music databases. To address this challenge, prior research has predominantly focused on a joint embedding of music audio and text, utilizing it to retrieve music tracks that exactly match descriptive queries related to musical attributes (i.e. genre, instrument) and contextual elements (i.e. mood, theme). However, users also articulate a need to explore music that shares similarities with their favorite tracks or artists, such as \textit{I need a similar track to Superstition by Stevie Wonder}. To address these concerns, this paper proposes an improved Text-to-Music Retrieval model, denoted as TTMR++, which utilizes rich text descriptions generated with a finetuned large language model and metadata. To accomplish this, we obtained various types of seed text from several existing music tag and caption datasets and a knowledge graph dataset of artists and tracks. The experimental results show the effectiveness of TTMR++ in comparison to state-of-the-art music-text joint embedding models through a comprehensive evaluation involving various musical text queries.

文本搜歌大模型音乐检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。