用扩散模型和可微混音台实现可控音乐混音,支持风格调节与精细参数优化。
Diff2Mix: Controllable Music Mixing via Diffusion Models and Differentiable Audio Effects

- 基于扩散模型与可微混音台,分两层控制混音风格和参数。
- 在客观与主观评测中均表现优秀,混音质量高且可编辑性强。
- 适合音乐制作人、音频工程师快速生成定制化混音结果。
自动音乐混音旨在将多轨录音融合为平衡且连贯的音乐作品。由于不同歌曲内容与混音工程师的主观偏好共同影响最终效果,实际系统需在保证高质量混音的同时支持可控的风格变化。然而,现有方法通常将自动混音与风格控制分开处理,难以实现高质量、可编辑且风格感知的统一系统。为此,本文提出 Diff2Mix,一种基于扩散模型与可微混音台的生成式自动混音系统。该系统提供两个层次的用户控制:参考音频实现整体制作风格控制,可微混音台则提供显式的音频效果参数,增强可解释性与细粒度优化能力。通过客观与主观评估,我们验证了系统在混音质量与控制能力方面的竞争力。代码与音频样例已公开于项目页面 https://zys711.github.io/Diff2Mix。
原文摘要 · Abstract (English)
Automatic music mixing aims to combine multitrack recordings into a balanced and coherent musical piece. Because the content of different songs and the subjective preferences of mixing engineers jointly shape the final outcome, a practical system should deliver well-balanced mixes while allowing for controllable stylistic variation. However, most existing methods treat automatic mixing and mixing style control as separate tasks, making it difficult for a single system to produce high-quality mixes while remaining editable and style-aware. To address this limitation, this paper presents Diff2Mix, a generative automatic mixing system based on diffusion models and a differentiable mixing console. This system offers two levels of optional user control: a reference audio enables overall production style control, and the differentiable mixing console provides explicit audio effects parameters for interpretability and fine-grained optimization. We demonstrate our system's competitive performance through both objective and subjective evaluations in terms of mixing quality and control ability. We provide code and audio samples at our project page https://zys711.github.io/Diff2Mix .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。