arXiv:2512.03479cs.CV2025-12被引 1

首个评估视频中物体动态演化的基准,揭示模型依赖语言先验而非细粒度物体变化。

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos

  • 构建跨视角的物体中心推理与时间证据定位联合评估框架
  • 13个主流多模态大模型在关键指标上均低于45%(mIoU)
  • 提供伪物体级监督方法,提升模型在规划任务中的迁移能力

程序性活动本质上由物体状态转变驱动,但现有教学视频基准仍以动作为中心,无法评估模型是否理解物体如何演化至任务完成。本文提出ProcObject-10K,首个同时评估物体中心推理与时间证据定位的教学视频基准,涵盖1,799段视频,10,522个开放式视频问答对,覆盖9个领域、137项任务及五类推理(前提条件、状态演变、反事实、错误分析、就绪判断)。对13个领先多模态大模型的基准测试显示显著的答案-定位差距:模型生成合理答案却无法准确定位支撑证据(mIoU < 45%),暴露其依赖语言先验而非细粒度物体动态。为进一步缩小差距,我们提供基于伪物体级监督与时空约束的微调基线。在ProcObject-10K上微调后的模型不仅性能提升,还能有效迁移到其他需要定位的视频问答与具身规划任务。数据集、标注与评估工具将公开发布,支持未来物体中心过程理解研究。

原文摘要 · Abstract (English)

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task completion. In this work, we introduce ProcObject-10K, the first benchmark that jointly evaluates object-centric reasoning and temporal evidence grounding in instructional videos, across both egocentric and exocentric views. It comprises 10,522 open-ended VideoQA pairs grounded in 1,799 video clips, spanning 137 tasks across 9 domains and five reasoning types covering preconditions, state evolution, counterfactuals, mistakes, and readiness. Benchmarking 13 leading MLLMs reveals a substantial answering-grounding gap: models produce plausible answers while failing to localize the supporting evidence (mIoU < 45%), exposing their reliance on linguistic priors rather than fine-grained object dynamics. As a step toward closing this gap, we further provide an object-centric supervised fine-tuning baseline with pseudo object-level supervision and spatial-temporal constraints. Models fine-tuned on ProcObject-10K not only improve on the benchmark itself, but also transfer effectively to other grounded VideoQA and embodied planning tasks. The dataset, annotations, and evaluation toolkit will be publicly released to support future research on object-centric procedural understanding.

视频理解物体中心推理评估具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。