首个面向音乐对话推荐的音频基准,推动跨模态推荐研究
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
- 构建真实用户对话与音乐音频的关联数据集
- 多模态测试显示当前模型在融合音频与文本时表现不佳
- 适合研究音乐推荐、跨模态学习的学者和工程师
对话式推荐因大语言模型迅速发展,但音乐推荐仍具独特挑战,需依赖音频内容而非仅靠文本或元数据。我们提出MusiCRS,首个面向音频中心的对话推荐基准,将Reddit上的真实用户对话与对应音乐曲目关联。该数据集包含477条高质量对话,涵盖古典、嘻哈、电子、金属、流行、独立、爵士等多元风格,涉及3,589个独特音乐实体,并通过YouTube链接实现音频定位。MusiCRS支持三种输入模态配置:仅音频、仅查询、音频+查询,可系统评估音频大模型、检索模型与传统方法。实验表明,当前系统在跨模态融合上表现有限,最优性能常出现在单模态设置而非多模态组合。这揭示了跨模态知识整合的根本局限——模型擅长对话语义理解,但在将抽象音乐概念与音频内容对齐时存在困难。为促进研究进展,我们开源MusiCRS数据集(https://huggingface.co/datasets/rohan2810/MusiCRS)、评估代码(https://github.com/rohan2810/musiCRS)及完整基线。
原文摘要 · Abstract (English)
Conversational recommendation has advanced rapidly with large language models (LLMs), yet music remains a uniquely challenging domain in which effective recommendations require reasoning over audio content beyond what text or metadata can capture. We present MusiCRS, the first benchmark for audio-centric conversational recommendation that links authentic user conversations from Reddit with corresponding tracks. MusiCRS includes 477 high-quality conversations spanning diverse genres (classical, hip-hop, electronic, metal, pop, indie, jazz), with 3,589 unique musical entities and audio grounding via YouTube links. MusiCRS supports evaluation under three input modality configurations: audio-only, query-only, and audio+query, allowing systematic comparison of audio-LLMs, retrieval models, and traditional approaches. Our experiments reveal that current systems struggle with cross-modal integration, with optimal performance frequently occurring in single-modality settings rather than multimodal configurations. This highlights fundamental limitations in cross-modal knowledge integration, as models excel at dialogue semantics but struggle when grounding abstract musical concepts in audio. To facilitate progress, we release the MusiCRS dataset (https://huggingface.co/datasets/rohan2810/MusiCRS), evaluation code (https://github.com/rohan2810/musiCRS), and comprehensive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。