arXiv:2511.04678cs.CV2025-11NeurIPS被引 5

让追踪系统在物体变形时仍能跟住并理解变化过程

Tracking and Understanding Object Transformations

  • 用图结构建模物体状态演化,零样本恢复变形后丢失目标
  • 在新数据集VOST-TAS上实现最佳追踪性能,准确率超基线12%
  • 适合研究视频理解、动态物体建模的学者和开发者

现实世界中的物体常经历状态变化,如苹果被切开或蝴蝶破茧而出。现有追踪方法因外观剧烈改变而容易丢失目标。为此,本文提出「跟踪任意状态」任务,旨在追踪物体在变形过程中的轨迹,并检测与描述状态变化,同时构建了新基准数据集VOST-TAS。针对该问题,提出TubeletGraph:一个零样本系统,通过语义与空间先验识别潜在遗漏轨迹,并判断是否合并;进而生成状态图以描述每阶段变化。该方法在复杂变形场景下保持领先追踪精度,在时间定位与语义推理方面表现优异。代码、结果与数据集已公开于https://tubelet-graph.github.io。

原文摘要 · Abstract (English)

Real-world objects frequently undergo state transformations. From an apple being cut into pieces to a butterfly emerging from its cocoon, tracking through these changes is important for understanding real-world objects and dynamics. However, existing methods often lose track of the target object after transformation, due to significant changes in object appearance. To address this limitation, we introduce the task of Track Any State: tracking objects through transformations while detecting and describing state changes, accompanied by a new benchmark dataset, VOST-TAS. To tackle this problem, we present TubeletGraph, a zero-shot system that recovers missing objects after transformation and maps out how object states are evolving over time. TubeletGraph first identifies potentially overlooked tracks, and determines whether they should be integrated based on semantic and proximity priors. Then, it reasons about the added tracks and generates a state graph describing each observed transformation. TubeletGraph achieves state-of-the-art tracking performance under transformations, while demonstrating deeper understanding of object transformations and promising capabilities in temporal grounding and semantic reasoning for complex object transformations. Code, additional results, and the benchmark dataset are available at https://tubelet-graph.github.io.

物体追踪状态演化视频理解零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。