多模态融合提升语音说话人分离鲁棒性,尤其在音视频质量差时仍表现优异。
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
- 融合音视频特征,用端到端二分类识别所有说话人活动。
- 在严重视频退化下性能仍接近最优音视频系统。
- 跨注意力机制适应不同说话人数,适合复杂场景应用。
本文提出一种质量感知的端到端音视频说话人分离框架,包含三项关键技术:首先,音视频模型同时输入音频与视觉特征,通过一系列二分类输出层同步识别所有说话人的活动;该端到端框架有效处理重叠语音,利用多模态信息准确区分语音与非语音段落。其次,采用质量感知的音视频融合结构,应对噪声、混响等音频退化,以及遮挡、离屏、检测不可靠等视频退化问题。最后,引入跨注意力机制于多说话人嵌入,使网络可适应不同数量说话人。实验结果表明,该方法在多种数据集上均表现出强鲁棒性,即使在视频质量严重退化情况下,性能也达到现有最佳音视频系统的水平。
原文摘要 · Abstract (English)
In this paper, we propose a quality-aware end-to-end audio-visual neural speaker diarization framework, which comprises three key techniques. First, our audio-visual model takes both audio and visual features as inputs, utilizing a series of binary classification output layers to simultaneously identify the activities of all speakers. This end-to-end framework is meticulously designed to effectively handle situations of overlapping speech, providing accurate discrimination between speech and non-speech segments through the utilization of multi-modal information. Next, we employ a quality-aware audio-visual fusion structure to address signal quality issues for both audio degradations, such as noise, reverberation and other distortions, and video degradations, such as occlusions, off-screen speakers, or unreliable detection. Finally, a cross attention mechanism applied to multi-speaker embedding empowers the network to handle scenarios with varying numbers of speakers. Our experimental results, obtained from various data sets, demonstrate the robustness of our proposed techniques in diverse acoustic environments. Even in scenarios with severely degraded video quality, our system attains performance levels comparable to the best available audio-visual systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。