arXiv:2605.31529cs.CV2026-05

构建动态赛场环境,评估视频智能的推理与决策能力。

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

论文配图:SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
图 1 · 摘自论文原文
  • 以团队运动为动态微世界,融合真实交互与可验证规则
  • 涵盖35000小时视频、1500万标注动作,支持多层级评估
  • 发现模型在自主推理任务中准确率仅5%,存在明显能力断层

真正的视频智能不仅需识别可见内容,更需理解事件成因、预测情境变化并制定下一步策略。我们将其归纳为从感知到因果推理、模拟再到战略规划的完整能力链,称为战略视频智能(SVI)。现有基准无法有效评估此能力体系:真实视频缺乏可验证的因果与策略真值,而合成环境又牺牲了真实多智能体系统的复杂性。为此,我们提出SVI-Bench,一个大规模基准,以团队运动为动态微世界,结合真实多智能体互动(10-22名参与者在对抗压力下协同决策)的复杂性与明确规则及确定结果的可验证性。该基准包含约35,000小时广播视频、1500万条标注动作、15,000小时专家解说、23,000份比赛报告及103,000条结构化统计数据,覆盖篮球、足球和冰球,全部通过数据引擎从原始比赛数据生成,形成密集交叉引用的知识库。评估分为9个任务,覆盖动态场景理解、因果推理、战略模拟与代理综合四个递进支柱。评估强大多模态与代理基线发现显著能力断层:模型在感知任务表现良好,细粒度动作问答达约74%准确率,但在后续认知层级迅速下降。代理任务最难:最强模型在需自主整合180万片段证据时,准确率仅5%。

原文摘要 · Abstract (English)

True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under different conditions, and deciding what to do next. We refer to this progression, from perception through causal reasoning and simulation to strategic planning, as Strategic Video Intelligence (SVI). No existing benchmark evaluates this capability stack: in-the-wild videos lack verifiable ground truth for causal and strategic questions, while synthetic environments sacrifice the complexity of real multi-agent systems. To bridge this gap, we introduce SVI-Bench, a large-scale benchmark that leverages team sports as a dynamic microworld, combining the complexity of real-world multi-agent interaction (10-22 agents making coordinated decisions under adversarial pressure) with the verifiability of explicit rules and definitive outcomes. SVI-Bench comprises approximately 35K hours of broadcast video, 15M annotated actions, 15K hours of expert commentary, 23K game reports, and 103K structured statistical records across basketball, soccer, and hockey, all constructed via a data engine that transforms raw game data into a dense, cross-referenced corpus. We organize evaluation into 9 tasks spanning a progressive four-pillar hierarchy: Dynamic Scene Understanding, Causal Reasoning, Strategic Simulation, and Agentic Synthesis. Evaluating strong multimodal and agentic baselines, we find a capability cliff: models perform competently on perceptual tasks, achieving approximately 74% on fine-grained action QA, but degrade sharply at each successive cognitive level. Agentic tasks prove hardest: the strongest model achieves only 5% accuracy when required to autonomously gather and integrate evidence across a corpus of 1.8M clips.

视频理解战略推理多智能体评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。