融合视听同步与人脸语音关联,提升第一视角视频中的说话人检测鲁棒性。
Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
- 通过加权平均融合两种模型输出,互补各自优劣。
- 在Ego4D数据集上分别达到70.2%和66.7%的mAP。
- 对低质量图像和遮挡场景有更好适应性,适合真实场景应用。
第一人称视角视频中的视听说话人检测(ASD)面临频繁遮挡、运动模糊和音频干扰等问题,导致唇动与语音的时间同步性难以辨识。传统基于同步的方法在干净条件下表现良好,但在第一人称视频中性能急剧下降;而基于人脸-语音关联(FVA)的方法虽能抵御瞬时视觉干扰,却在多人重叠说话或前端分割错误时表现不佳。本文提出一种简单有效的集成方法,通过加权平均融合依赖同步与不依赖同步的模型输出,充分利用互补线索,无需复杂融合结构。同时优化了FVA组件的预处理流程以增强集成效果。在Ego4D-AVD验证集上的实验表明,该集成方法使用TalkNet和Light-ASD作为主干网络时,分别获得70.2%和66.7%的均值平均精度(mAP)。按人脸图像质量与话语遮蔽程度分层的定性分析进一步验证了两种组件的互补优势。
原文摘要 · Abstract (English)
Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech. Traditional synchronisation-based systems perform well under clean conditions but degrade sharply in first-person recordings. Conversely, face-voice association (FVA)-based methods forgo synchronisation modelling in favour of cross-modal biometric matching, exhibiting robustness to transient visual corruption but suffering when overlapping speech or front-end segmentation errors occur. In this paper, a simple yet effective ensemble approach is proposed to fuse synchronisation-dependent and synchronisation-agnostic model outputs via weighted averaging, thereby harnessing complementary cues without introducing complex fusion architectures. A refined preprocessing pipeline for the FVA-based component is also introduced to optimise ensemble integration. Experiments on the Ego4D-AVD validation set demonstrate that the ensemble attains 70.2% and 66.7% mean Average Precision (mAP) with TalkNet and Light-ASD backbones, respectively. A qualitative analysis stratified by face image quality and utterance masking prevalence further substantiates the complementary strengths of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。