构建19万条真实对话音乐推荐数据集,兼顾规模与真实性。
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation

- 从Reddit社区提取19万组对话对,关联Deezer音乐元数据。
- 保留原始对话真实性,同时提供可复现的改写版本。
- 适合研究真实场景下基于内容的对话推荐系统。
当前对话式音乐推荐研究面临真实语料规模有限与合成语料失真的两难。本文提出Reddit2Deezer,一个基于19万组{线程, 叶子评论}对的真实世界对话音乐推荐数据集。数据以两种形式发布:保留原始对话真实性的原始版本,以及为提升长期可复现性而进行改写的版本。每条音乐实体均链接至Deezer标识符,可直接获取音频试听及丰富元数据(如流派标签、热度、BPM)。人工验证确认了对话质量、物品定位准确性和改写效果。数据集已公开于https://huggingface.co/datasets/McAuley-Lab/Reddit2Deezer。
原文摘要 · Abstract (English)
Conversational music recommendation (CMR) research currently faces a tradeoff between authentic dialogue corpora that are limited in scale and synthesized corpora that scale up but whose conversations are artificially constructed rather than naturally observed. In this paper, we introduce Reddit2Deezer, a reality-grounded CMR resource derived from 190k unique {thread, leaf-comment} pairs. We release the resource in two versions: a raw version that preserves authenticity, and a paraphrased version that maximizes long-term reproducibility. Each musical entity is linked to a Deezer identifier, which provides straightforward access to audio previews and rich metadata (e.g., genre tags, popularity, BPM), opening the door to future research on content-grounded conversational recommendation. A human validation confirms the quality of the dialogues, item grounding, and paraphrases. The dataset is available at https://huggingface.co/datasets/McAuley-Lab/Reddit2Deezer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。