通过唇部关键点提升说话人检测在复杂环境下的鲁棒性
LASER: Lip Landmark Assisted Speaker Detection for Robustness
- 用唇部关键点引导模型关注语音相关区域,增强注意力聚焦
- 在高噪声场景下,相比LoCoNet和TalkNet,mAP分别提升3.3和4.3点
- 无需测试时依赖唇部检测器,对遮挡和低分辨率有更强适应性
主动说话人检测(ASD)旨在识别复杂视觉场景中的发言者。尽管人类自然依赖唇动与音频同步,现有模型在唇动与音频不同步时常误判非发言状态。为此,我们提出唇部关键点辅助的鲁棒说话人检测方法(LASER),在训练中显式引入唇部关键点,引导模型关注语音相关区域。对于人脸轨迹,LASER提取视觉特征并将其2D唇部关键点编码为密集图。为应对低分辨率或遮挡等失败情况,我们设计了辅助一致性损失,对齐含唇信息与仅人脸预测,使测试时无需唇部检测器。LASER在域内与域外基准上均优于当前最优模型。为进一步评估真实场景下的鲁棒性,我们构建了LASER-bench数据集,包含带不同背景噪声的现代视频片段。在高噪声子集上,LASER相较LoCoNet和TalkNet分别提升mAP 3.3和4.3点,展现出对真实声学挑战的强大抗性。
原文摘要 · Abstract (English)
Active Speaker Detection (ASD) aims to identify who is speaking in complex visual scenes. While humans naturally rely on lip-audio synchronization, existing ASD models often misclassify non-speaking instances when lip movements and audio are unsynchronized. To address this, we propose Lip landmark Assisted Speaker dEtection for Robustness (LASER), which explicitly incorporates lip landmarks during training to guide the model's attention to speech-relevant regions. Given a face track, LASER extracts visual features and encodes 2D lip landmarks into dense maps. To handle failure cases such as low resolution or occlusion, we introduce an auxiliary consistency loss that aligns lip-aware and face-only predictions, removing the need for landmark detectors at test time. LASER outperforms state-of-the-art models across both in-domain and out-of-domain benchmarks. To further evaluate robustness in realistic conditions, we introduce LASER-bench, a curated dataset of modern video clips with varying levels of background noise. On the high-noise subset, LASER improves mAP by 3.3 and 4.3 points over LoCoNet and TalkNet, respectively, demonstrating strong resilience to real-world acoustic challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。