arXiv:2603.07758cs.CV2026-03中稿 · CVPR被引 1

通过背景锚点实现视频中目标长时间追踪,解决遮挡与重入难题。

AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos

  • 利用静态背景构建锚点库,生成持久语义记忆
  • 重入时提升捕捉率10.3%,延迟降低24.2%
  • 适合长时视频目标定位,尤其处理遮挡与离场场景

固定视角视频中的长期语言引导定位极具挑战:目标可能长时间被遮挡或离开画面后重新出现,而逐帧引用方法因重识别(ReID)不可靠导致漂移。AR2-4FV利用背景稳定性实现长期定位。离线构建从静态背景结构中提取的锚点库;推理时,文本查询与该库对齐生成锚点图,作为目标缺席时的持久语义记忆。基于锚点的重入先验加速目标重捕获,轻量级ReID-Gating机制利用锚点帧位移线索维持身份连续性。系统无需假设目标首帧可见,也未显式建模外观变化,即可预测每帧边界框。相较于最佳基线,该方法实现+10.3%重捕获率(RCR)提升和-24.2%重捕获延迟(RCL)降低,消融实验进一步验证了锚点图、重入先验与ReID-Gating的有效性。

原文摘要 · Abstract (English)

Long-term language-guided referring in fixed-view videos is challenging: the referent may be occluded or leave the scene for long intervals and later re-enter, while framewise referring pipelines drift as re-identification (ReID) becomes unreliable. AR2-4FV leverages background stability for long-term referring. An offline Anchor Bank is distilled from static background structures; at inference, the text query is aligned with this bank to produce an Anchor Map that serves as persistent semantic memory when the referent is absent. An anchor-based re-entry prior accelerates re-capture upon return, and a lightweight ReID-Gating mechanism maintains identity continuity using displacement cues in the anchor frame. The system predicts per-frame bounding boxes without assuming the target is visible in the first frame or explicitly modeling appearance variations. AR2-4FV achieves +10.3% Re-Capture Rate (RCR) improvement and -24.2% Re-Capture Latency (RCL) reduction over the best baseline, and ablation studies further confirm the benefits of the Anchor Map, re-entry prior, and ReID-Gating.

视频定位目标追踪长时记忆锚点机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。