arXiv:2506.17886cs.SDeess.AS2025-06中稿 · ISMIR 2025被引 4

用扩散模型生成可调控的音乐检索查询,提升交互灵活性。

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

  • 通过扩散模型在优化后的潜在空间生成查询,实现可控生成。
  • 相比对比教师模型,检索性能提升,支持非联合训练编码器的音频空间检索。
  • 支持负向提示、DDIM反演等后处理操作,增强用户交互控制能力。

多模态对比模型在文本-音频检索和零样本设置中表现强劲,但联合嵌入空间的优化仍是活跃研究方向。现有系统较少关注用户可控性与交互性。在文本-音乐检索中,自由文本的模糊性导致多对多映射,常引发结果僵化或不满足需求。本文提出生成式扩散检索器(GDR),利用扩散模型在检索优化的潜在空间中生成查询,通过负向提示、去噪扩散隐式模型(DDIM)反演等生成工具实现可控性,开辟了检索控制新路径。GDR在性能上超越对比教师模型,并支持使用非联合训练编码器的纯音频潜在空间检索。最后,实验表明其能有效实现检索行为的后处理操控,显著增强文本-音乐检索任务中的交互控制能力。

原文摘要 · Abstract (English)

Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems controllable and interactive for users. In text-music retrieval, the ambiguity of freeform language creates a many-to-many mapping, often resulting in inflexible or unsatisfying results. We introduce Generative Diffusion Retriever (GDR), a novel framework that leverages diffusion models to generate queries in a retrieval-optimized latent space. This enables controllability through generative tools such as negative prompting and denoising diffusion implicit models (DDIM) inversion, opening a new direction in retrieval control. GDR improves retrieval performance over contrastive teacher models and supports retrieval in audio-only latent spaces using non-jointly trained encoders. Finally, we demonstrate that GDR enables effective post-hoc manipulation of retrieval behavior, enhancing interactive control for text-music retrieval tasks.

文本音乐检索扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。