arXiv:2507.12972eess.AScs.SD2025-07

音频视觉融合模型可灵活分离任意数量说话人,提升嘈杂环境下的语音分离效果。

AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning

  • 多尺度编码与并行架构联合优化计数与分离任务
  • 在多个数据集上达到当前最佳性能,支持任意数量说话人分离
  • 适合真实复杂环境中需要自适应说话人数量的语音分离场景

从包含任意数量说话人的混合信号中分离目标语音是一项挑战性任务。现有方法虽在分离性能和抗噪能力方面表现优异,但普遍依赖对说话人数量的先验知识。针对未知说话人数量场景的研究有限,其在真实声学环境中的泛化能力显著受限。为此,本文提出AVFSNet——一种结合多尺度编码与并行架构的音视频语音分离模型,联合优化说话人计数与多人语音分离任务。模型可并行独立分离每个说话人,并通过引入视觉信息增强对环境噪声的适应能力。全面实验表明,AVFSNet在多个评估指标上达到当前最优结果,且在多种数据集上表现出色。

原文摘要 · Abstract (English)

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge of speaker counts in mixtures. The limited research addressing unknown speaker quantity scenarios exhibits significantly constrained generalization capabilities in real acoustic environments. To overcome these challenges, this paper proposes AVFSNet -- an audio-visual speech separation model integrating multi-scale encoding and parallel architecture -- jointly optimized for speaker counting and multi-speaker separation tasks. The model independently separates each speaker in parallel while enhancing environmental noise adaptability through visual information integration. Comprehensive experimental evaluations demonstrate that AVFSNet achieves state-of-the-art results across multiple evaluation metrics and delivers outstanding performance on diverse datasets.

语音分离音视频融合多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。