arXiv:2604.21873cs.CV2026-04被引 1

提出首个物理信号对齐的视频理解评测基准,精准定位事件时空位置。

Grounding Video Reasoning in Physical Signals

论文配图:Grounding Video Reasoning in Physical Signals
图 1 · 摘自论文原文
  • 构建四类视频源、六领域物理任务的可复现评测体系
  • 物理提示表现最优,空间定位仍是最弱环节
  • 强调需报告提示类型与扰动敏感性,避免误判模型能力

视频物理理解不仅需要正确命名事件,还需准确定位事件的时间与空间位置。现有模型可能仅依赖文本规律回答倒水、滑动或碰撞问题,却无法精确定位。本文提出一个基于物理信号的视频理解基准,扩展V-STaR的what--when--where评估框架,涵盖四类视频源(SSV2、YouCook2、HoloAssist、Roundabout-TAU)、六类物理领域、三类提示(物理型、vstar_like、neutral_rstr)及四种输入条件(原始、打乱、剔除、帧掩码)。共包含1,560个基础视频片段,每段经统一事件记录生成,三类提示均共享时间与空间标签,非物理提示使用基于同一记录生成的确定性语义目标。结果显示:整体上物理提示表现最佳,vstar_like为清晰的非物理语义对照,neutral_rstr作为更难的模板化控制;提示鲁棒性呈选择性而非普适性,扰动增益集中在原模型表现弱的场景,空间定位在所有设置中均为最弱。结论:视频问答评测应同时报告物理对齐、提示感知与扰动敏感性诊断,而非仅依赖聚合准确率。

原文摘要 · Abstract (English)

Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textual regularities while still failing to localize the event in time or space. We introduce a grounded benchmark for physical video understanding that extends the what--when--where evaluation structure of V-STaR to four video sources, six physics domains, three prompt families (physics, vstar_like, and neutral_rstr), and four input conditions (original, shuffled, ablated, and frame-masked). The benchmark contains 1,560 base video clips from SSV2, YouCook2, HoloAssist, and Roundabout-TAU. Each clip is first converted into a shared grounded event record, and the three query families are derived from that record. Temporal and spatial targets are shared across prompt families, while the non-physics families use deterministic family-appropriate semantic a_what targets derived from the same record. Across models and prompt families, physics remains the strongest regime overall, vstar_like is the clearest non-physics semantic comparison, and neutral_rstr behaves as a harder templated control. Prompt-family robustness is selective rather than universal, perturbation gains cluster in weak original cases, and spatial grounding is the weakest across settings. These results suggest that video Q&A reasoning benchmarks shall report physically grounded, prompt-aware, and perturbation-aware diagnostics alongside aggregate accuracy.

视频理解物理建模评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。