提出视频中物体状态变化的空间渐进分割任务,精准定位变化发生的区域和速度。
SPOC: Spatially-Progressing Object State Change Segmentation in Video
- 基于视觉语言模型生成伪标签,引入状态变化动态约束。
- 在两个数据集上验证了新任务的挑战性与模型定位精度。
- 适用于机器人追踪动作进展,推动更敏感的状态感知研究。
视频中的物体状态变化揭示了人类和智能体行为的关键线索。然而,现有方法仅能定位物体处于初始状态(如奶酪块)与完成状态变化(如磨碎的奶酪)的时间点,无法提供变化发生的空间位置信息。本文提出空间渐进式物体状态变化分割任务,目标是像素级分割出物体中可操作区域与已转变区域。实验表明,当前先进的视觉语言模型(VLM)与视频分割方法在此任务上表现不佳,凸显其难度与新颖性。为此,我们设计了一种基于VLM的伪标签方法、状态变化动态约束,并构建了基于真实网络视频的WhereToChange基准数据集。在两个数据集上的实验验证了该任务的挑战性及所提模型在精确识别变化位置与速度上的潜力。进一步展示了其在追踪活动进展方面的应用价值,有助于提升机器人对动态场景的理解能力。整体而言,本工作将空间物体状态变化分割确立为视频理解的新前沿任务,挑战现有最先进方法,激励社区构建更鲁棒、更敏感的状态变化表征。
原文摘要 · Abstract (English)
Object state changes in video reveal critical cues about human and agent activity. However, existing methods are limited to temporal localization of when the object is in its initial state (e.g., cheese block) versus when it has completed a state change (e.g., grated cheese), offering no insight into where the change is unfolding. We propose to deepen the problem by introducing the spatially-progressing object state change segmentation task. The goal is to segment at the pixel-level those regions of an object that are actionable and those that are transformed. We show that state-of-the-art VLMs and video segmentation methods struggle at this task, underscoring its difficulty and novelty. As an initial baseline, we design a VLM-based pseudo-labeling approach, state-change dynamics constraints, and a novel WhereToChange benchmark built on in-the-wild Internet videos. Experiments on two datasets validate both the challenge of the new task as well as the promise of our model for localizing exactly where and how fast objects are changing in video. We further demonstrate useful implications for tracking activity progress to benefit robotic agents. Overall, our work positions spatial OSC segmentation as a new frontier task for video understanding: one that challenges current SOTA methods and invites the community to build more robust, state-change-sensitive representations. Project page: https://vision.cs.utexas.edu/projects/spoc-spatially-progressing-osc
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。