评测视频大模型的物理异常推理能力,区分物体自身属性与环境关系异常。
Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs
- 基于本体因果分类法,系统划分物体属性与物理关系异常。
- 20个视频大模型中识别任务得分75-88,解释任务F1低于50。
- 适合研究视觉推理、模型可解释性及多模态理解的学者使用。
我们提出Video-HOCA,一个用于评估视频大模型物理异常推理能力的诊断基准。该基准采用本体因果分类法,区分实体自身属性或能力的违反与实体间及与环境间物理关系的违反。数据集包含超过1,400个生成和真实视频,以及3,470个问答对,所有标签和参考答案均经人工验证。评估涵盖四个推理层级:合理性检查、异常归因、细粒度识别和开放式物理解释。在20个指令模式的视频大模型中,识别任务得分集中在75-88之间,而异常归因任务的宏平均F1值普遍低于50。我们还发现,本体因果差距依赖于任务类型和模型配置,且思维模式提升不能仅由采样或输出预算解释。通过标注一致性、第四任务人类评分与评分者间一致性、替代指标及时间/解码控制,验证了评估流程的有效性,并限定了基准支持的结论范围。
原文摘要 · Abstract (English)
We introduce Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos. Video-HOCA uses an Ontological-Causal taxonomy to distinguish violations of an entity's own properties or capabilities from violations of physical relations among entities and the environment. It contains more than 1,400 generated and real-world videos and 3,470 question-answer pairs, with human verification of labels and reference answers. The benchmark evaluates four levels of reasoning: plausibility checking, anomaly attribution, fine-grained recognition, and open-ended physical explanation. Across 20 Instruct-mode Video-LLMs, we find that recognition outpaces explanation: Task I scores cluster at 75-88, while Task II macro-F1 stays mostly below 50. We also find that the Ontological-Causal gap depends on the task and model configuration, and that Thinking-mode gains are not explained by sampling or output budget alone. Annotation agreement, Task-IV human-judge and judge-judge checks, alternative metrics, and temporal/decoding controls validate the evaluation pipeline and bound the claims supported by the benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。