arXiv:2412.04746cs.SDcs.IR2024-12被引 4

用扩散模型生成多样音乐搜索方向,更懂用户模糊偏好。

Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance

  • 用轻量扩散模型从用户描述生成多个潜在音乐搜索方向
  • 在检索和排序指标上优于传统方法,推荐更相关且多样
  • 支持文本或图像引导,适合需要灵活探索的音乐发现场景

现代音乐检索系统通常依赖固定的用户偏好表示,难以捕捉用户多样化和不确定的检索需求。为解决此问题,我们提出 Diff4Steer,一种新颖的生成式检索框架,利用轻量级扩散模型从用户查询中合成多样化的种子嵌入,代表潜在的音乐探索方向。与将用户查询映射到嵌入空间单一位置的确定性方法不同,Diff4Steer为音频目标模态提供了统计先验,有效捕捉用户偏好的不确定性和多维度特征。此外,Diff4Steer可由图像或文本输入引导,结合最近邻搜索实现更灵活、可控的音乐发现。实验表明,该框架在检索与排序指标上均优于确定性回归方法及基于大语言模型的生成式检索基线,证明其在捕捉用户偏好方面的有效性,带来更丰富且相关的推荐结果。听觉示例见 tinyurl.com/diff4steer。

原文摘要 · Abstract (English)

Modern music retrieval systems often rely on fixed representations of user preferences, limiting their ability to capture users' diverse and uncertain retrieval needs. To address this limitation, we introduce Diff4Steer, a novel generative retrieval framework that employs lightweight diffusion models to synthesize diverse seed embeddings from user queries that represent potential directions for music exploration. Unlike deterministic methods that map user query to a single point in embedding space, Diff4Steer provides a statistical prior on the target modality (audio) for retrieval, effectively capturing the uncertainty and multi-faceted nature of user preferences. Furthermore, Diff4Steer can be steered by image or text inputs, enabling more flexible and controllable music discovery combined with nearest neighbor search. Our framework outperforms deterministic regression methods and LLM-based generative retrieval baseline in terms of retrieval and ranking metrics, demonstrating its effectiveness in capturing user preferences, leading to more diverse and relevant recommendations. Listening examples are available at tinyurl.com/diff4steer.

音乐生成扩散模型检索增强语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。