arXiv:2605.09259cs.SDcs.AI2026-05

直接从混音中实现多乐器音色迁移,避免分离误差。

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

论文配图:Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems
图 1 · 摘自论文原文
  • 用联合扩散模型同步迁移多个音轨的音色
  • 在SATB合唱数据集上优于单音轨基线模型
  • 适合音乐制作与跨乐器音色重构场景

音色迁移旨在改变音乐录音的音色特征,同时保留原始旋律与节奏。尽管单乐器音色迁移已取得显著进展,现有多乐器方法依赖于先分离再迁移的流水线,会传播源分离误差并导致各音轨合成音色不连贯。本文提出MixtureTT,据我们所知首个可直接从多声部混合信号中实现灵活逐音轨音色迁移的系统。给定一个混合音频及每个目标声部的独立音色参考,MixtureTT通过共享扩散过程联合将所有音轨迁移至指定乐器。通过建模各音轨内容间及跨音轨谐波间的依赖关系,所提出的联合音轨扩散变压器消除了级联分离误差,推理成本降低为音轨数倍,且输出更连贯。尽管在更难的输入条件下运行,SATB合唱数据集上的评估显示,MixtureTT在客观与主观指标上均优于单音轨基线模型,证实了专用多乐器音色迁移对简单分离-迁移流水线的必要性。结果表明,跨音轨建模对混合级音色迁移至关重要,所提联合设置始终优于等效的单音轨消融实验。

原文摘要 · Abstract (English)

Timbre transfer aims to modify the timbral identity of a musical recording while preserving the original melody and rhythm. While single-instrument timbre transfer has made substantial progress, existing approaches to multi-instrument settings rely on separate-then-transfer pipelines that propagate source separation artifacts and produce incoherent synthesized timbres across stems. This paper proposes MixtureTT, to the best of our knowledge the first system for flexible per-stem timbre transfer directly from a polyphonic mixture. Given a mixture and a separate timbre reference for each target voice, MixtureTT jointly transfers all stems to the specified instruments through a shared diffusion process. Modeling the dependencies across the per-stem content and cross-stem harmonic, the proposed joint stem diffusion transformer eliminates cascaded separation error, reduces inference cost by a factor equal to the number of stems, and yields more coherent multi-stem outputs. Despite operating under a strictly harder input condition, evaluations on the SATB choral dataset show that MixtureTT outperforms single-instrument baselines on both objective and subjective metrics demonstrating the necessity of dedicated multi-instrument timbre transfer over the naive separate-then-transfer pipelines. As a result, this work confirms that the cross-stem modeling is essential for mixture-level timbre transfer as the proposed joint setting consistently exceeds an equivalent single-stem ablation.

音色迁移扩散模型多音轨处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。