不依赖干音参考,通过对比学习捕捉音频效果的相对差异。
Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning

- 用双分支编码器和交叉注意力提取音频间的效果变换关系。
- 在MUSDB18上四种乐器均超越现有方法,性能领先。
- 无需干音数据,适合真实音乐制作场景中的效果建模。
音频效果(Fx)表征学习在智能音乐制作中至关重要,包括自动混音与效果风格迁移。现有方法通常依赖干音或近乎干音的参考信号进行建模,但实际录音不可避免地受麦克风、声学环境及前置处理影响,完全未处理的音频极少存在。我们提出,相比绝对效果编码,音频间的相对效果距离更符合真实生产需求。为此,我们设计了RelFx——一种无需干音参考即可从通用音频集合中学习相对效果变换的对比学习框架。该方法采用双分支Siamese编码器,结合交叉注意力与差分门控融合,从参考片段与内容相关的效果处理片段中推断共享效果变换。此外,我们提出反称融合变体,实现双向效果编码:交换输入顺序可得近似符号相反的嵌入,此性质此前未被探索。所提方法摆脱对干音多轨数据集的依赖,支持在含效果的音频上训练。在标准的Fx-Encoder++ MUSDB18评估协议下,效果风格迁移任务表现达到当前最优,四类乐器均持续优于现有方法。
原文摘要 · Abstract (English)
Audio effects (Fx) representation learning plays a key role in intelligent music production, including automatic mixing and Fx style transfer. Existing methods typically rely on dry or nearly dry references for effect modeling, yet truly unprocessed audio is rarely available in practice, as real recordings inevitably reflect the microphone, room acoustics, and preceding signal processing. Instead of pursuing absolute effect encodings, we argue that the relative effect distance between audio signals is more meaningful for real-world music production. Motivated by this, we propose RelFx, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training. Our approach uses a dual-branch Siamese encoder equipped with cross-attention and differential gating fusion to infer the shared effect transformation from a reference clip and an effect-processed, content-related clip. We further propose an antisymmetric fusion variant for bidirectional effect encoding, such that swapping the input order directly produces a nearly sign-reversed embedding, a property not explored in earlier work. Moreover, our dry-reference-free formulation eliminates the reliance on dry multitrack datasets and enables training on effect-bearing audio. Experiments on Fx style transfer demonstrate state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol, consistently outperforming existing approaches across all four instrument categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。