构建音乐查询意图数据集,揭示大模型对隐含意图的捕捉短板
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
- 人工标注2291条Reddit音乐请求,区分七类描述符的正/负/参照偏好
- 大模型能识别明确描述,但对依赖上下文的意图理解能力不足
- 为音乐推荐系统优化提供可复用的意图解析基准
尽管带有用户查询标注的音乐描述符数据集日益增多,但很少关注这些描述符背后的用户意图,而意图对于有效满足用户需求至关重要。本文提出 MusicRecoIntent,一个包含 2,291 条 Reddit 音乐请求的手动标注语料库,对七类音乐描述符标注其在查询中扮演的正向、负向或参照性偏好角色。随后,我们研究了大型语言模型(LLMs)提取这些音乐描述符的可靠性,发现它们能够捕捉显式描述符,但在处理依赖上下文的描述符时表现不佳。该工作可作为细粒度用户意图建模的基准,并为改进基于 LLM 的音乐理解系统提供洞见。
原文摘要 · Abstract (English)
Although annotated music descriptor datasets for user queries are increasingly common, few consider the user's intent behind these descriptors, which is essential for effectively meeting their needs. We introduce MusicRecoIntent, a manually annotated corpus of 2,291 Reddit music requests, labeling musical descriptors across seven categories with positive, negative, or referential preference-bearing roles. We then investigate how reliably large language models (LLMs) can extract these music descriptors, finding that they do capture explicit descriptors but struggle with context-dependent ones. This work can further serve as a benchmark for fine-grained modeling of user intent and for gaining insights into improving LLM-based music understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。