arXiv:2608.23330cs.CV2026-08TPAMI被引 2

让机器看懂视频中人的意图,突破视觉识别局限。

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

论文配图:IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
图 1 · 摘自论文原文
  • 构建三类认知上下文增强视频意图理解
  • 在对比测试集上性能下降率降低40%以上
  • 推理过程可解释,适合需要透明AI的场景

视频理解需超越视觉事实识别,深入理解人类行为背后的意图(常被称为社会智能的‘暗物质’)。为弥合视觉观察与意图推理之间的差距,我们提出新任务IntentQA,并构建大规模视频问答数据集。为避免标准指标因数据偏差而高估模型能力,我们引入五组对比数据集和‘对比性能下降’评估指标。提出X-CaVIR框架,融合三类认知上下文:情境上下文(通过跨模态视频查询语言模块)、对比上下文(对比学习模块)和常识上下文(常识推理模块)。关键创新在于采用透明流程整合LLM,结合视频描述与VQA输出,不仅提升性能,还使推理过程可解释。实验表明,该框架优于现有基线,在对比扰动下表现更稳定。

原文摘要 · Abstract (English)

Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR (eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the opacity of traditional black-box models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.

视频理解意图识别可解释AI常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。