为日常对话推荐合适的背景音乐,构建了首个标准化评测基准。
DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues
- 基于1200段对话和4个候选音乐片段,构建带人类偏好排序的评测集。
- 现有模型最高仅35%命中率,远低于人类水平。
- 适合研究对话理解与音乐生成融合的AI开发者和媒体工程师。
为自然人机对话选择合适的背景音乐(BGM)是媒体和交互系统中的常见需求。本文提出对话条件下的BGM推荐任务,要求模型从无明确音乐描述的多轮对话中,选出不干扰、贴合语境的音乐。为此,我们构建了DialBGM基准,包含1,200段开放域日常对话,每段对话配有4个候选音乐片段,并由人工标注偏好排序。评分依据包括语境相关性、非侵入性和一致性。我们评估了多种开源与专有模型,涵盖音频-语言模型及多模态大模型,结果表明当前模型在选择排名第一音乐片段时,最高命中率仅为35%,显著落后于人类判断。DialBGM为开发具有话语意识的BGM选择方法提供了标准化评测平台,可用于评估检索与生成两类模型。
原文摘要 · Abstract (English)
Selecting an appropriate background music (BGM) that supports natural human conversation is a common production step in media and interactive systems. In this paper, we introduce dialogue-conditioned BGM recommendation, where a model should select non-intrusive, fitting music for a multi-turn conversation that often contains no music descriptors. To study this novel problem, we present DialBGM, a benchmark of 1,200 open-domain daily dialogues, each paired with four candidate music clips and annotated with human preference rankings. Rankings are determined by background suitability criteria, including contextual relevance, non-intrusiveness, and consistency. We evaluate a wide range of open-source and proprietary models, including audio-language models and multimodal LLMs, and show that current models fall far short of human judgments; no model exceeds 35% Hit@1 when selecting the top-ranked clip. DialBGM provides a standardized benchmark for developing discourse-aware methods for BGM selection and for evaluating both retrieval-based and generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。