arXiv:2507.05822cs.CV2025-07被引 1

用大模型知识增强视觉模型,让机器能推理视频中发生了什么及接下来会怎样。

Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models

  • 融合视觉基础模型与大语言模型,实现视觉与常识推理的协同
  • 在多个基准上达到顶尖性能,零样本泛化能力显著
  • 适合需要理解视频因果关系与未来预测的智能系统研究者

当前视频理解模型擅长识别‘发生了什么’,但在因果推理和未来预测等高层认知任务上表现不足,根源在于缺乏常识性世界知识。为此,我们提出一种新框架,将强大的视觉基础模型(VFM)与作为知识驱动推理核心的大语言模型(LLM)协同融合。关键技术是受Q-Former启发的融合模块,可将复杂的时空与物体中心视觉特征提炼为简洁、语言对齐的表示,使LLM能基于直接视觉证据进行推断。模型采用两阶段训练策略:先在大规模视频-文本数据上进行对齐预训练,再在精心设计的数据集上进行指令微调以激发高级推理与预测能力。大量实验表明,该模型在多个挑战性基准上达到最先进水平,尤其展现出对未见推理任务的出色零样本泛化能力;深入的消融研究验证了各组件的关键作用。本工作推动机器感知从简单识别迈向真正的认知理解,为机器人、人机交互等领域更智能的AI系统开辟道路。

原文摘要 · Abstract (English)

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To bridge this cognitive gap, we propose a novel framework that synergistically fuses a powerful Vision Foundation Model (VFM) for deep visual perception with a Large Language Model (LLM) serving as a knowledge-driven reasoning core. Our key technical innovation is a sophisticated fusion module, inspired by the Q-Former architecture, which distills complex spatiotemporal and object-centric visual features into a concise, language-aligned representation. This enables the LLM to effectively ground its inferential processes in direct visual evidence. The model is trained via a two-stage strategy, beginning with large-scale alignment pre-training on video-text data, followed by targeted instruction fine-tuning on a curated dataset designed to elicit advanced reasoning and prediction skills. Extensive experiments demonstrate that our model achieves state-of-the-art performance on multiple challenging benchmarks. Notably, it exhibits remarkable zero-shot generalization to unseen reasoning tasks, and our in-depth ablation studies validate the critical contribution of each architectural component. This work pushes the boundary of machine perception from simple recognition towards genuine cognitive understanding, paving the way for more intelligent and capable AI systems in robotics, human-computer interaction, and beyond.

视频理解大模型融合因果推理零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。