arXiv:2603.14733cs.CV2026-03被引 1

构建多视频理解新框架与评测集,提升跨视频推理能力

A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding

  • 设计技能增强型智能体框架,支持多视频迭代推理
  • 在4255个视频上构建1442道题的评测集,覆盖11类任务
  • 适合研究多模态推理、视频理解与智能体系统的研究者

多模态大模型在单视频理解上表现强劲,但在跨视频推理方面仍受限。现有方法通常将多个视频拼接后直接输入,导致训练与推理不一致、帧压缩引发信息损失,且缺乏显式的跨视频协同机制。当前多视频评测主要聚焦事件级对比,忽视身份级匹配、细粒度区分和结构化多步推理。为此,我们提出MVX-Bench,一个涵盖11类经典计算机视觉任务的统一多视频问答评测基准,包含来自多样化真实数据集的4,255个视频和1,442个问题。同时,我们提出SAMA——一种技能增强型智能体框架,融合视觉工具、任务特定技能和冲突感知验证机制,实现迭代式结构化推理。实验表明,SAMA在MVX-Bench上优于主流开源基线和GPT,消融实验证明技能设计与冲突解决机制的有效性。

原文摘要 · Abstract (English)

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.

多视频理解智能体框架评测基准跨视频推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。