用结构化场景图增强视觉语言模型的视频问答推理能力
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- 引入场景图作为中间表征,补足视觉语言模型的时空推理短板
- 在三个基准上提升因果与时间推理性能,优于已有基线方法
- 适合关注视频理解可解释性与混合智能系统的研究人员
视频问答(VQA)需要模型对视频中的空间、时间与因果线索进行推理。近期的视觉语言模型(VLMs)虽表现强劲,但常依赖浅层相关性,导致时间定位能力弱且可解释性差。本文研究将结构化场景图(SGs)作为中间接地信号用于VQA。SGs提供物体-关系的结构化表示,可补充VLM的整体推理。我们提出SG-VLM,一种模块化框架,通过提示与视觉定位将冻结的VLM与场景图接地结合。在三个基准(NExT-QA、iVQA、ActivityNet-QA)和多个VLM(QwenVL、InternVL)上,SG-VLM提升了因果与时间推理能力,优于先前基线,尽管相对于强基线的提升有限。这些发现揭示了符号接地的潜力与当前局限,为未来融合VLM与符号方法的视频理解提供了指导。
原文摘要 · Abstract (English)
Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often rely on shallow correlations, leading to weak temporal grounding and limited interpretability. We study symbolic scene graphs (SGs) as intermediate grounding signals for VQA. SGs provide structured object-relation representations that complement VLMs holistic reasoning. We introduce SG-VLM, a modular framework that integrates frozen VLMs with scene graph grounding via prompting and visual localization. Across three benchmarks (NExT-QA, iVQA, ActivityNet-QA) and multiple VLMs (QwenVL, InternVL), SG-VLM improves causal and temporal reasoning and outperforms prior baselines, though gains over strong VLMs are limited. These findings highlight both the promise and current limitations of symbolic grounding, and offer guidance for future hybrid VLM-symbolic approaches in video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。