arXiv:2510.21485cs.SDeess.AS2025-10被引 2

FlexIO可灵活处理任意麦克风和说话人数量的语音分离与增强。

FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement

  • 用提示向量条件化分离,支持任意数量说话人
  • 多通道输入通过无阵列依赖机制统一处理
  • 在1-5麦克风、1-3说话人下均表现稳健

语音分离与增强(SSE)在受控环境下已取得显著进展,如固定说话人数和阵列配置。为构建通用SSE系统,单通道系统已扩展至处理可变说话人数,而多通道系统也发展出适配不同阵列配置的能力。然而,这些方向长期独立演进。本文提出一种灵活输入输出的SSE系统——FlexIO,采用每说话人一个提示向量进行条件化分离,实现任意数量说话人的分离。多通道混合信号通过无阵列依赖的通道通信机制与提示向量联合处理。实验表明,FlexIO在1至5个麦克风、1至3个说话人的多种条件下均有效,且在真实数据集CHiME-4上验证了鲁棒性。

原文摘要 · Abstract (English)

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data.

语音分离多通道灵活架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。