让音乐自动适配不同音箱,提升听感一致性。
Device-Guided Music Transfer
- 用视觉语言模型分析音箱频响曲线,提取设备特征。
- 在自建数据集上微调,实现少量样本即可适配新音箱。
- 适合做音箱风格增强或音质优化的开发者使用。
设备引导的音乐迁移旨在为无法接触真实设备的用户,在未见设备上实现播放适配。现有方法多聚焦于修改音色、节奏、和声或配器以模仿特定流派或歌手,却忽视了播放设备(如扬声器)的多样硬件特性。为此,我们提出DeMT,将扬声器的频率响应曲线作为线图输入,通过视觉-语言模型提取设备嵌入,并利用特征逐维线性调制(feature-wise linear modulation)控制混合变换器。在自建数据集上微调后,DeMT实现了有效的音箱风格迁移与对未见设备的鲁棒少样本适应,支持音箱风格增强与音质提升等应用。
原文摘要 · Abstract (English)
Device-guided music transfer adapts playback across unseen devices for users who lack them. Existing methods mainly focus on modifying the timbre, rhythm, harmony, or instrumentation to mimic genres or artists, overlooking the diverse hardware properties of the playback device (i.e., speaker). Therefore, we propose DeMT, which processes a speaker's frequency response curve as a line graph using a vision-language model to extract device embeddings. These embeddings then condition a hybrid transformer via feature-wise linear modulation. Fine-tuned on a self-collected dataset, DeMT enables effective speaker-style transfer and robust few-shot adaptation for unseen devices, supporting applications like device-style augmentation and quality enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。