用轻量迭代融合提升音视频语音分离效果
Audio-Visual Speech Separation via Bottleneck Iterative Network
- 通过瓶颈融合令牌实现轻量级迭代特征优化
- 在多个数据集上性能超越主流模型,训练与推理提速超50%
- 适合需要高效音视频分离的实时应用场景
利用非听觉线索可显著提升语音分离模型性能。现有方法通常使用深层模态专用网络提取单模态特征,但存在计算成本高或容量不足的问题。本文提出一种名为瓶颈迭代网络(BIN)的迭代表示精炼方法:通过轻量级融合模块反复迭代,利用融合令牌对融合表示进行瓶颈压缩,从而在不显著增加模型规模的前提下提升模型容量,并平衡性能与训练成本。我们在噪声环境下的音视频语音分离任务中测试了BIN,在NTCD-TIMIT和LRS3+WHAM!数据集上均实现了比现有最优模型更高的SI-SDRi指标,同时几乎所有设置下训练和GPU推理时间均减少超过50%。
原文摘要 · Abstract (English)
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。