arXiv:2604.05079cs.CV2026-04被引 4

让AI像人一样看懂视频剧情,提升问答准确率。

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

  • 用多个智能体协作构建视频故事线,逐步优化理解过程。
  • 在多个VideoQA数据集上表现优于现有方法,最高提升6.2%。
  • 适合需要可解释视频理解的场景,如教育、医疗分析。

视频问答(VideoQA)是一项挑战性任务,需融合空间、时间与语义信息以捕捉视频序列的复杂动态。尽管近期已有多种视频理解方法,但多数仍依赖定位相关帧回答问题,而非像人类一样通过推理演变的故事线进行理解。人类天然通过连贯故事线解读视频,这对做出鲁棒且上下文相关的预测至关重要。为此,我们提出SVAgent,一种基于故事线引导的跨模态多智能体框架用于视频问答。故事线智能体基于修正建议智能体分析历史失败所提示的帧,逐步构建叙事表示。同时,跨模态决策智能体在故事线引导下,独立从视觉和文本模态预测答案;其输出由元智能体评估,以对齐跨模态预测并增强推理鲁棒性与答案一致性。实验表明,通过模拟人类式的故事线推理,SVAgent在多个基准上实现更优性能与更高可解释性。

原文摘要 · Abstract (English)

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches for video understanding, most existing methods still rely on locating relevant frames to answer questions rather than reasoning through the evolving storyline as humans do. Humans naturally interpret videos through coherent storylines, an ability that is crucial for making robust and contextually grounded predictions. To address this gap, we propose SVAgent, a storyline-guided cross-modal multi-agent framework for VideoQA. The storyline agent progressively constructs a narrative representation based on frames suggested by a refinement suggestion agent that analyzes historical failures. In addition, cross-modal decision agents independently predict answers from visual and textual modalities under the guidance of the evolving storyline. Their outputs are then evaluated by a meta-agent to align cross-modal predictions and enhance reasoning robustness and answer consistency. Experimental results demonstrate that SVAgent achieves superior performance and interpretability by emulating human-like storyline reasoning in video understanding.

视频理解多智能体故事线推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。