融合音视频信息的注意力机制,显著降低语音分离错误率。
Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
- 用跨模态与自注意力捕捉音视频交互和上下文关系。
- 训练中结合伪标签优化,使错误率降至8.18%。
- 适合需要高精度音视频说话人分离的应用场景。
本文介绍了为2025年多模态信息语音处理(MISP)挑战赛任务1开发的系统。提出CASA-Net,一种面向端到端音视频说话人分离(AVSD)系统的嵌入融合方法。CASA-Net引入跨注意力(CA)模块以有效捕获音视频信号间的跨模态交互,并采用自注意力(SA)模块学习音视频帧之间的上下文关系。为进一步提升性能,采用融合伪标签精炼与重训练的训练策略,提高时间戳预测准确率。此外,应用中值滤波与重叠平均作为后处理技术,消除异常值并平滑预测标签。系统在评测集上实现8.18%的说话人分离错误率(DER),相比基线15.52%相对提升47.3%。
原文摘要 · Abstract (English)
This paper presents the system developed for Task 1 of the Multi-modal Information-based Speech Processing (MISP) 2025 Challenge. We introduce CASA-Net, an embedding fusion method designed for end-to-end audio-visual speaker diarization (AVSD) systems. CASA-Net incorporates a cross-attention (CA) module to effectively capture cross-modal interactions in audio-visual signals and employs a self-attention (SA) module to learn contextual relationships among audio-visual frames. To further enhance performance, we adopt a training strategy that integrates pseudo-label refinement and retraining, improving the accuracy of timestamp predictions. Additionally, median filtering and overlap averaging are applied as post-processing techniques to eliminate outliers and smooth prediction labels. Our system achieved a diarization error rate (DER) of 8.18% on the evaluation set, representing a relative improvement of 47.3% over the baseline DER of 15.52%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。