用预训练音频与视觉对齐提升第一视角说话人检测,效果领先。
EgoVIS@CVPR: PAIR-Net: Enhancing Egocentric Speaker Detection via Pretrained Audio-Visual Fusion and Alignment Loss
- 融合冻结的Whisper音频编码器与微调的AV-HuBERT视觉模型。
- 引入跨模态对齐损失,在真实场景下达76.6% mAP,领先8.2%-12.9%。
- 适合第一视角视频中语音不可见或视角不稳的场景应用。
第一视角视频中的主动说话人检测面临视角不稳、运动模糊及声音源离屏等挑战,传统视觉主导方法性能显著下降。我们提出PAIR-Net(基于预训练音视频融合与正则化网络),将部分冻结的Whisper音频编码器与微调的AV-HuBERT视觉主干结合,实现跨模态信息有效融合。为缓解模态不平衡问题,设计了跨模态对齐损失,使音视频表征同步,促进各模态稳定收敛。无需依赖多说话人上下文或理想正面视角,PAIR-Net在Ego4D ASD基准上达到76.6% mAP,较LoCoNet和STHG分别提升8.2%和12.9%。结果表明,预训练音频先验与基于对齐的融合对真实世界第一视角条件下的鲁棒说话人检测具有重要价值。
原文摘要 · Abstract (English)
Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We introduce PAIR-Net (Pretrained Audio-Visual Integration with Regularization Network), an effective model that integrates a partially frozen Whisper audio encoder with a fine-tuned AV-HuBERT visual backbone to robustly fuse cross-modal cues. To counteract modality imbalance, we introduce an inter-modal alignment loss that synchronizes audio and visual representations, enabling more consistent convergence across modalities. Without relying on multi-speaker context or ideal frontal views, PAIR-Net achieves state-of-the-art performance on the Ego4D ASD benchmark with 76.6% mAP, surpassing LoCoNet and STHG by 8.2% and 12.9% mAP, respectively. Our results highlight the value of pretrained audio priors and alignment-based fusion for robust ASD under real-world egocentric conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。