用人脸与声音的生物特征关联,解决第一视角录音中的说话人检测难题。
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
- 基于视觉质量动态加权帧,融合人脸身份信息增强鲁棒性。
- 参数量少于传统方法,仍达到相近甚至更优的检测性能。
- 适合摄像头遮挡、模糊或噪声大的第一视角场景使用。
音频视频说话人检测(ASD)通常依赖声学与视觉语音信号的时间同步性建模。然而在第一视角录像中,遮挡、运动模糊和恶劣声学条件会削弱基于同步的方法效果。本文提出一种新框架,仅利用跨模态人脸-声音关联来判断说话人活动。将现有脸-音关联模型与基于Transformer的编码器结合,通过动态加权每帧视觉质量以聚合面部身份信息。该系统再与前端语音段分割方法集成,形成完整ASD系统。实验表明,所提系统SL-ASD在参数量显著减少的情况下,性能可媲美甚至超越参数密集的同步方法,验证了在挑战性第一视角场景中,以灵活的生物特征关联替代严格音视频同步建模的可行性。
原文摘要 · Abstract (English)
Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is compromised by occlusions, motion blur, and adverse acoustic conditions. In this work, a novel framework is proposed that exclusively leverages cross-modal face-voice associations to determine speaker activity. An existing face-voice association model is integrated with a transformer-based encoder that aggregates facial identity information by dynamically weighting each frame based on its visual quality. This system is then coupled with a front-end utterance segmentation method, producing a complete ASD system. This work demonstrates that the proposed system, Self-Lifting for audiovisual active speaker detection(SL-ASD), achieves performance comparable to, and in certain cases exceeding, that of parameter-intensive synchronisation-based approaches with significantly fewer learnable parameters, thereby validating the feasibility of substituting strict audiovisual synchronisation modelling with flexible biometric associations in challenging egocentric scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。