arXiv:2609.07226cs.SDcs.LG2026-09

通过MIMO架构实现稳定迭代语音分离,保持音源相位与音色一致性。

Iterative Audio Separation with Mixture Consistency via MIMO Model Extension

论文配图:Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
图 1 · 摘自论文原文
  • 将分离模型扩展为多输入多输出结构,支持迭代优化
  • 在多个基准上显著提升分离性能,尤其在相位保真度上
  • 适合需要精确音色和相位信息的应用,如语音增强

本文提出一种通用框架,通过将源分离模型扩展为多输入多输出(MIMO)配置,实现稳定且有效的迭代音频分离,并保持混合信号的一致性。在音频分离领域,混合一致性对需要准确目标音源相位和音色信息的应用至关重要。尽管扩散模型等迭代方法在语音增强或用户引导的音源分离任务中表现出优越的听觉效果,但多数现有方法仍局限于单步分离,采用单输入单输出(SISO)或单输入多输出(SIMO)结构,因为混合一致性分离通常被视为具有唯一解的回归问题。通过将这些架构扩展至MIMO配置,我们实现了无需牺牲架构优势或混合一致性特性的迭代预测。我们对结合判别器及扩展为生成模型进行了全面消融研究。实验结果表明,将该框架应用于最先进分离模型时,性能有显著提升。

原文摘要 · Abstract (English)

This paper proposes a general framework for stable and effective iterative audio separation with mixture consistency by extending source separation models to a multi-input multi-output (MIMO) configuration. In the field of audio separation, mixture consistency is an essential property for many applications that require accurate phase and timbral information of target sources. While iterative approaches such as diffusion models achieve perceptually superior results in speech enhancement or user-guided target source separation tasks, most existing methods focus on single-step separation with a single-input single-output (SISO) or single-input multi-output (SIMO) configuration through architectural improvements, since mixture-consistent audio separation is generally regarded as a regression problem that admits a unique solution. By extending these architectures to a MIMO configuration, we introduce iterative prediction without compromising the architectural advantages or the characteristics of mixture consistency. We conduct a comprehensive ablation study of combining the framework with discriminators and extending it to a generative model. Experimental results demonstrate significant performance improvements when applying the proposed framework to state-of-the-art separation models.

音频分离MIMO迭代生成相位一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。