用词相对坐标定位物体状态变化,提升视觉追踪精度。
Déjà Cue: Localizing States in Object Histories via Vocabulary-Relative Coordinates

- 构建词汇相对坐标系,通过对比描述差异定位状态区间。
- 在78个物体轨迹上,R@1提升至20.5%,Top-1 tIoU达21.5%。
- 无需训练,可直接用于已有模型的状态检索,适合视觉追踪研究者。
追踪通过视觉变化关联同一物体的观测,但无法判断物体是空还是满、完整还是被切割。本文提出身份条件下的状态时刻检索:给定一个追踪物体的历史和多个状态描述,定位每个状态成立的时间区间。传统方法依赖图像-文本相似度独立评估描述,但因每一帧都显示同一目标,共享对象兼容性会掩盖状态证据。通过引入备选描述作为参照,能有效提取状态差异。我们提出Déjà Cue,一种无需训练的框架,将备选描述转化为词汇相对坐标系。该方法从每条描述中减去其状态平衡质心,校准帧得分,并使用冻结编码器扫描连续可见片段中的多种时长。在78个VOST物体轨迹上,固定时间扫描仅改变查询参考,R@1在tIoU 0.5下从10.3%提升至20.5%,Top-1 tIoU从16.0%升至21.5%。候选排名分析表明,词汇相对查询能更优地排序有效区间。因此,相关状态描述可作为对象特定、查询时的坐标系统,用于解读冻结的视觉表征。
原文摘要 · Abstract (English)
Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut. We formulate identity-conditioned state-moment retrieval: given a tracked-object history and alternative state descriptions, localize an interval in which each described state holds. Absolute image-text similarity scores descriptions independently; because every visible frame depicts the same target, shared object compatibility can obscure the state evidence needed to identify the target interval. The alternatives provide the missing reference: evidence for one state should be measured against the others. We introduce Déjà Cue, a training-free framework that turns these alternatives into a vocabulary-relative coordinate system. It subtracts their state-balanced centroid from each description, calibrates frame scores, and scans multiple durations within contiguous visible runs using a frozen encoder. On 78 VOST histories, holding the temporal scan fixed and changing only the query reference nearly doubles R@1 at tIoU 0.5 from 10.3\% to 20.5\% and raises Top-1 tIoU from 16.0\% to 21.5\%. Candidate-rank analyses show that vocabulary-relative queries rank useful intervals higher within the same candidate set. Related state descriptions can therefore serve as an object-specific, query-time coordinate system for reading frozen visual representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。