把混音变成逐轨叠加,更像人操作。
Rethinking Automatic Music Mixing as Sequential Stem Blending

- 按顺序逐轨混入已有音轨,模拟人工混音流程。
- 在多个基准上表现优于传统并行方法,提升音质一致性。
- 适合需要灵活处理多轨的音乐制作场景。
自动混音通常通过并行架构一次性处理所有音轨。受人类混音师逐轨操作启发,本文提出将自动混音重构为序列化音轨融合任务:每次将一个音轨融入不断增长的子混音中。为此,我们训练了一个基于子混音上下文的隐空间流匹配模型,可处理任意数量输入音轨。为训练模型,引入一种基于退化的数据合成策略,从现有多轨和源分离数据集中模拟真实音轨融合场景。在音轨融合与自动混音基准测试中,实验结果证明该方法有效性。音频示例见配套演示页。
原文摘要 · Abstract (English)
Automatic music mixing, the task of automatically combining individual audio tracks into a cohesive mixture, is typically addressed by parallelized architectures that process all input tracks in a single pass. In this work, inspired by how human mix engineers process stems one at a time, we propose a paradigm shift and ask whether automatic music mixing can be reformulated as a sequential stem blending task, where each stem is blended into a growing submix. Specifically, we train a latent flow matching model conditioned on the submix context, enabling sequential processing of an arbitrary number of input tracks. To train the model, we introduce a degradation-based data synthesis strategy that simulates realistic stem blending scenarios from existing multitrack and source separation datasets. Experimental results on both stem blending and automatic music mixing benchmarks demonstrate the effectiveness of the proposed approach. We provide audio examples on the accompanying demo page\footnote{https://sequential-mixing-demo.vercel.app/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。