用大模型生成音乐描述,提升跨模态相似度检索效果
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
- 用大模型生成音乐文本描述,弥补高质量图文数据稀缺
- 在华为音乐平台实测,检索准确率显著优于现有方法
- 适合音乐推荐系统、跨模态搜索方向的研究者参考
音乐相似性检索是流媒体平台管理与探索大规模内容的基础。本文提出一种新型跨模态对比学习框架,利用开放式的文本描述引导音乐相似性建模,克服传统单模态方法在捕捉复杂音乐关系上的局限。为解决高质量文-音配对数据稀缺问题,本文引入双源数据获取方法,结合在线爬取与大模型提示生成,通过精心设计的提示词利用大模型的音乐知识生成语境丰富的描述。大量实验表明,该框架在客观指标、主观评估及华为音乐平台真实A/B测试中均显著优于现有基准。
原文摘要 · Abstract (English)
Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs' comprehensive music knowledge to generate contextually rich descriptions. Exten1sive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。