通过反事实对比构建,提升第一视角视频问答的多事件与交互理解能力。
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
- 设计双模态反事实样本生成:文本用事件改写,视觉用核心交互挖掘。
- 在EgoTaskQA上达到52.51%(正常)和46.04%(间接)准确率,领先当前水平。
- 适合关注第一视角视频理解、交互识别与模型鲁棒性研究的读者。
第一人称视频问答(Egocentric VideoQA)在第一视角视频理解中具有重要意义,旨在基于第一人称视角视频回答问题。尽管现有方法通过预训练-微调范式取得进展,但忽视了第一人称视角的独特挑战,如多事件理解与手物交互识别。为此,我们提出双模态反事实对比构建(DMC³)框架,包含一个第一人称视频问答基线、反事实样本构建模块及涉及反事实样本的对比优化模块。具体而言,我们首先设计反事实样本构建模块,通过事件描述改写生成文本模态的正负样本,通过核心交互挖掘生成视觉模态的正负样本。随后将这些样本与原始样本一同输入基线模型。最后,在反事实样本参与的对比优化模块中,应用对比损失使原始样本特征与正样本特征距离最小化,与负样本特征距离最大化。实验表明,本方法在EgoTaskQA的'normal'和'indirect' splits上分别达到52.51%和46.04%的准确率,在QAEGO4D上达到13.2%的准确率,均达到当前最优性能。
原文摘要 · Abstract (English)
Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos. Although existing methods have made progress through the paradigm of pre-training and fine-tuning, they ignore the unique challenges posed by the first-person perspective, such as understanding multiple events and recognizing hand-object interactions. To deal with these challenges, we propose a Dual-Modal Counterfactual Contrastive Construction (DMC$^3$) framework, which contains an egocentric videoqa baseline, a counterfactual sample construction module and a counterfactual sample-involved contrastive optimization. Specifically, We first develop a counterfactual sample construction module to generate positive and negative samples for textual and visual modalities through event description paraphrasing and core interaction mining, respectively. Then, We feed these samples together with the original samples into the baseline. Finally, in the counterfactual sample-involved contrastive optimization module, we apply contrastive loss to minimize the distance between the original sample features and the positive sample features, while maximizing the distance from the negative samples. Experiments show that our method achieve 52.51\% and 46.04\% on the \textit{normal} and \textit{indirect} splits of EgoTaskQA, and 13.2\% on QAEGO4D, both reaching the state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。