解析自动混音中两阶段设计的有效性,揭示分步处理的优势。
Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System

- 通过分任务分析,拆解局部平衡与全局协调
- 错误分组导致下游性能明显下降,音量关系影响较小
- 两阶段系统显著优于单阶段,适合复杂音乐混音
自动混音将多轨录音转化为听觉连贯、平衡且美学一致的混音作品。在实际制作中,由于轨道数量多、乐器多样且轨道间依赖性强,该任务极具挑战。两阶段系统通过将组内处理与组间混音分离来应对复杂性,但其性能提升究竟源于更强的组件模型,还是显式的任务分解尚不明确。本文通过三个受控听觉实验,分析自动混音的子任务机制:考察全混音模型是否可迁移至组内混音、下游模型能否补偿分组与音量误差、以及两阶段分解是否提升整体混音质量。在三个密集的流行与摇滚片段上测试发现,不同模型间的迁移效果存在差异;不当分组会引发明显下游退化,而音量关系改变的影响较弱且依赖模型。两种两阶段变体均显著优于对应单阶段基线。结果支持将局部平衡与全局混音协调显式分离作为自动混音的设计原则。代码与音频示例已公开。
原文摘要 · Abstract (English)
Automatic mixing transforms multitrack recordings into perceptually coherent, balanced, and aesthetically consistent mixes. In real-world production, this task is challenging due to large track counts, diverse instrumentation, and strong inter-track dependencies. Two-stage systems address this complexity by separating intra-group processing from inter-group mixing, yet it remains unclear whether their gains arise from stronger component models or from explicit task decomposition. We present a subtask-oriented analysis of automatic mixing through three controlled listening experiments. We investigate whether full-mix models transfer to intra-group mixing, whether downstream models compensate for grouping and loudness errors, and whether two-stage decomposition improves full-mix quality. Across three dense pop and rock excerpts, transfer differs between the evaluated models; inappropriate grouping causes clear downstream degradation, while altered loudness relationships have weaker and model-dependent effects. Both two-stage variants significantly outperform their corresponding single-stage baselines. These findings support explicit separation of local balance and global mix coordination as a useful design principle for automatic mixing. Code and audio examples are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。