arXiv:2409.07747cs.CVcs.AI2024-09被引 2

用图模型提升视频问答中多物体事件的因果推理能力

Multi-object event graph representation learning for Video Question Answering

  • 构建多物体事件图,通过对比学习对齐问题与图结构
  • 在NExT-QA和TGIF-QA-R上准确率提升2.2%,因果类问题高2.8%
  • 适合需要复杂时空关系推理的视频理解任务

视频问答(VideoQA)旨在根据给定视频回答问题,系统需理解视频中物体之间的时空关系以进行因果与时间推理。现有基于Transformer的方法虽能建模单个物体运动,但在涉及多个物体的复杂场景(如“男孩将球投进篮筐”)中表现不佳。为此,本文提出一种对比语言事件图表示学习方法CLanG,通过多层GNN-聚类模块进行对抗性图表示学习,实现问题文本与相关多物体事件图之间的对比学习。该方法在两个挑战性数据集NExT-QA和TGIF-QA-R上超越强基线,最高提升2.2%准确率;尤其在因果与时间类问题上优于基线2.8%,展现出对多物体事件推理的强大能力。

原文摘要 · Abstract (English)

Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships among objects extracted from videos to perform causal and temporal reasoning. While prior works have focused on modeling individual object movements using transformer-based methods, they falter when capturing complex scenarios involving multiple objects (e.g., "a boy is throwing a ball in a hoop"). We propose a contrastive language event graph representation learning method called CLanG to address this limitation. Aiming to capture event representations associated with multiple objects, our method employs a multi-layer GNN-cluster module for adversarial graph representation learning, enabling contrastive learning between the question text and its relevant multi-object event graph. Our method outperforms a strong baseline, achieving up to 2.2% higher accuracy on two challenging VideoQA datasets, NExT-QA and TGIF-QA-R. In particular, it is 2.8% better than baselines in handling causal and temporal questions, highlighting its strength in reasoning multiple object-based events.

视频问答事件图多物体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。