用生成模型模拟混音效果嵌入,实现多风格自动混音。
Automatic Music Mixing using a Generative Model of Effect Embeddings
- 通过生成轨道嵌入控制效果器,建模专业混音的分布。
- 在多种音乐类型上接近人类混音水平,优于现有方法。
- 支持无标签干声与湿声训练,可处理任意数量音轨。
音乐混音是将多个音轨融合为整体的过程,具有主观性,同一输入可能对应多个有效混音结果。现有自动混音系统将其视为确定性回归问题,忽略了这种解的多样性。本文提出MEGAMI(Multitrack Embedding Generative Auto MIxing),一种生成式框架,用于建模给定未处理音轨时专业混音的条件分布。MEGAMI采用不依赖音轨的效应处理器,以每轨生成的嵌入作为条件;通过排列等变架构处理任意数量的未标注音轨;并利用领域自适应技术,实现对干声和湿声录音的联合训练。客观评估使用分布度量显示性能持续优于现有方法;听觉测试表明,在多种音乐类型下表现接近人类水平。
原文摘要 · Abstract (English)
Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic regression problem, thus ignoring this multiplicity of solutions. Here we introduce MEGAMI (Multitrack Embedding Generative Auto MIxing), a generative framework that models the conditional distribution of professional mixes given unprocessed tracks. MEGAMI uses a track-agnostic effects processor conditioned on per-track generated embeddings, handles arbitrary unlabeled tracks through a permutation-equivariant architecture, and enables training on both dry and wet recordings via domain adaptation. Our objective evaluation using distributional metrics shows consistent improvements over existing methods, while listening tests indicate performances approaching human-level quality across diverse musical genres.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。