让音乐变 mood 可控,换情绪不改风格。
Controllable Embedding Transformation for Mood-Guided Music Retrieval
- 用情绪标签引导嵌入转换,实现仅改情绪不改风格。
- 在两个数据集上成功改变情绪,同时保留原曲流派与配器。
- 适合需要精准情绪匹配的音乐推荐场景。
音乐表征是现代推荐系统的核心,支撑歌单生成、相似性搜索与个性化发现。然而现有嵌入大多无法单独调控某一音乐属性,例如仅改变歌曲情绪而不影响其流派或配器。本文提出一种基于嵌入变换的情绪引导音乐检索框架,旨在保留种子音频的其他特征的同时,根据指定情绪标签将嵌入映射至目标嵌入。由于无法直接修改原始音频的情绪,我们引入采样机制,通过检索代理目标来平衡多样性与与种子的相似性。训练一个轻量级转换模型,并设计新颖的联合目标函数以同时优化变换效果与信息保留。在两个数据集上的大量实验表明,该方法在情绪转换上表现优异,同时显著优于无需训练的基线,在流派和配器保留方面也大幅提升,确立了可控嵌入变换在个性化音乐检索中的潜力。
原文摘要 · Abstract (English)
Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer little control for adjusting a single musical attribute, e.g., changing only the mood of a track while preserving its genre or instrumentation. In this work, we address the problem of controllable music retrieval through embedding-based transformation, where the objective is to retrieve songs that remain similar to a seed track but are modified along one chosen dimension. We propose a novel framework for mood-guided music embedding transformation, which learns a mapping from a seed audio embedding to a target embedding guided by mood labels, while preserving other musical attributes. Because mood cannot be directly altered in the seed audio, we introduce a sampling mechanism that retrieves proxy targets to balance diversity with similarity to the seed. We train a lightweight translation model using this sampling strategy and introduce a novel joint objective that encourages transformation and information preservation. Extensive experiments on two datasets show strong mood transformation performance while retaining genre and instrumentation far better than training-free baselines, establishing controllable embedding transformation as a promising paradigm for personalized music retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。