EgoAdapt提升第一人称语音对话检测在缺失视觉信息时的鲁棒性。
EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
- 融合头部朝向与唇动信息,捕捉言语与非言语线索。
- 在Ego4D数据集上实现62.01%准确率,优于当前最佳方法4.96%。
- 动态感知模态缺失,适合真实复杂场景下的交互研究。
TMT(与我交谈)任务是理解人类社交互动的关键,旨在确定谁正与佩戴摄像头者对话。传统模型在真实场景中常因视觉数据缺失、忽略头部朝向和背景噪声而表现不佳。本文提出EgoAdapt,一种针对模态缺失下第一人称TMT Speaker Detection的自适应框架。其包含三个模块:(1) 视觉说话者目标识别(VSTR)模块,利用头部朝向作为非语言线索、唇动作为语言线索,综合解析言语与非言语信号;(2) 并行共享权重音频(PSA)编码器,提升嘈杂环境下的音频特征提取能力;(3) 视觉模态缺失感知(VMMA)模块,实时估计每帧各模态是否存在,动态调整系统响应。在Ego4D数据集的TMT基准上全面评估,EgoAdapt达到67.39%的mAP和62.01%的Acc,相比当前最优方法分别提升1.56%和4.96%。
原文摘要 · Abstract (English)
TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to missing visual data, neglecting the role of head orientation, and background noise. This study addresses these limitations by introducing EgoAdapt, an adaptive framework designed for robust egocentric "Talking to Me" speaker detection under missing modalities. Specifically, EgoAdapt incorporates three key modules: (1) a Visual Speaker Target Recognition (VSTR) module that captures head orientation as a non-verbal cue and lip movement as a verbal cue, allowing a comprehensive interpretation of both verbal and non-verbal signals to address TTM, setting it apart from tasks focused solely on detecting speaking status; (2) a Parallel Shared-weight Audio (PSA) encoder for enhanced audio feature extraction in noisy environments; and (3) a Visual Modality Missing Awareness (VMMA) module that estimates the presence or absence of each modality at each frame to adjust the system response dynamically.Comprehensive evaluations on the TTM benchmark of the Ego4D dataset demonstrate that EgoAdapt achieves a mean Average Precision (mAP) of 67.39% and an Accuracy (Acc) of 62.01%, significantly outperforming the state-of-the-art method by 4.96% in Accuracy and 1.56% in mAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。