从混音中分离出每种乐器的音频效果特征,提升智能音乐制作能力。
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
- 通过对比学习与查询机制,将混音效果嵌入转换为乐器级效果表示。
- 在多种乐器上均优于传统混音级方法,实现精准效果匹配。
- 适合需要精细音频处理的音乐制作工具开发者使用。
通用音频表征已在多种音乐信息检索任务中表现优异,但在智能音乐制作中的应用受限于对音频效果(Fx)理解不足。现有方法多聚焦于混音级效果分析,难以满足自动混音等需乐器级效果理解的任务需求。本文提出Fx-Encoder++,一种新型模型,可从音乐混音中提取乐器级音频效果表征。该方法采用对比学习框架,引入“提取器”机制,通过提供乐器查询(音频或文本),将混音级音频效果嵌入转化为乐器级效果嵌入。我们在检索与音频效果参数匹配任务上评估模型性能,涵盖多种乐器。结果表明,Fx-Encoder++不仅优于以往混音级方法,还首次展现出乐器级效果表示提取能力,填补了智能音乐制作系统的关键能力空白。
原文摘要 · Abstract (English)
General-purpose audio representations have proven effective across diverse music information retrieval applications, yet their utility in intelligent music production remains limited by insufficient understanding of audio effects (Fx). Although previous approaches have emphasized audio effects analysis at the mixture level, this focus falls short for tasks demanding instrument-wise audio effects understanding, such as automatic mixing. In this work, we present Fx-Encoder++, a novel model designed to extract instrument-wise audio effects representations from music mixtures. Our approach leverages a contrastive learning framework and introduces an "extractor" mechanism that, when provided with instrument queries (audio or text), transforms mixture-level audio effects embeddings into instrument-wise audio effects embeddings. We evaluated our model across retrieval and audio effects parameter matching tasks, testing its performance across a diverse range of instruments. The results demonstrate that Fx-Encoder++ outperforms previous approaches at mixture level and show a novel ability to extract effects representation instrument-wise, addressing a critical capability gap in intelligent music production systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。