arXiv:2510.11715cs.CV2025-10被引 5

用视频扩散模型零样本追踪目标点,靠颜色标记自动跟蹤運動軌跡。

Point Prompting: Counterfactual Tracking with Video Diffusion Models

  • 用彩色標記提示擴散模型,讓它在生成過程中跟蹤點的移動。
  • 在遮擋下仍能持續追蹤,性能接近專用自監督模型。
  • 僅需初始圖像作負面提示,無需訓練,適合快速實驗。

追蹤器與視頻生成器解決密切相關的問題:前者分析運動,後者合成運動。我們證明,預訓練的視頻擴散模型可透過簡單提示實現零樣本點追蹤——只需指示模型在運動過程中可視化標記點。在查詢點放置一個獨特顏色的標記,然後從中間噪聲水平重新生成視頻,使標記沿著時間傳播,追蹤點的軌跡。為確保標記在這種反事實生成中保持可見(儘管自然視頻中幾乎不會出現此類標記),我們使用未編輯的初始幀作為負面提示。通過多種圖像條件化視頻擴散模型的實驗,我們發現這些「衍生」追蹤結果超越了以往零樣本方法,且在遮擋情況下仍具魯棒性,通常表現出與專用自監督模型競爭的水平。

原文摘要 · Abstract (English)

Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffusion models to perform zero-shot point tracking by simply prompting them to visually mark points as they move over time. We place a distinctively colored marker at the query point, then regenerate the rest of the video from an intermediate noise level. This propagates the marker across frames, tracing the point's trajectory. To ensure that the marker remains visible in this counterfactual generation, despite such markers being unlikely in natural videos, we use the unedited initial frame as a negative prompt. Through experiments with multiple image-conditioned video diffusion models, we find that these "emergent" tracks outperform those of prior zero-shot methods and persist through occlusions, often obtaining performance that is competitive with specialized self-supervised models.

視頻追蹤擴散模型零樣本反事實生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。