arXiv:2505.24404cs.CV2025-05被引 1

融合音视频线索提升第一人称社交互动检测精度

PCIE_Interaction Solution for Ego4D Social Interaction Challenge

  • 通过人脸质量增强与集成学习提升注视检测性能
  • 音视频联合建模,按视觉质量加权融合,实现0.71 mAP
  • 适合关注第一人称视频社交分析的研究者参考

本报告介绍我们团队在CVPR 2025 Ego4D社交互动挑战赛中针对Looking At Me(LAM)和Talking To Me(TTM)任务提出的PCIE_Interaction解决方案。LAM任务仅依赖人脸裁剪序列,而TTM任务结合说话人面部图像与同步音频片段。针对LAM,采用人脸质量增强与集成方法;针对TTM,通过视觉质量得分加权融合音视频线索。最终在排行榜上分别取得0.81和0.71的平均精度(mAP)。代码已公开于https://github.com/KanokphanL/PCIE_Ego4D_Social_Interaction。

原文摘要 · Abstract (English)

This report presents our team's PCIE_Interaction solution for the Ego4D Social Interaction Challenge at CVPR 2025, addressing both Looking At Me (LAM) and Talking To Me (TTM) tasks. The challenge requires accurate detection of social interactions between subjects and the camera wearer, with LAM relying exclusively on face crop sequences and TTM combining speaker face crops with synchronized audio segments. In the LAM track, we employ face quality enhancement and ensemble methods. For the TTM task, we extend visual interaction analysis by fusing audio and visual cues, weighted by a visual quality score. Our approach achieved 0.81 and 0.71 mean average precision (mAP) on the LAM and TTM challenges leader board. Code is available at https://github.com/KanokphanL/PCIE_Ego4D_Social_Interaction

社交互动音视频融合第一人称视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。