arXiv:2410.10053cs.CV2024-10NeurIPS被引 5

用扩散模型插值实现更高效精准的目标跟踪

DINTR: Tracking via Diffusion-based Interpolation

  • 基于扩散模型的插值机制,直接重建视频帧
  • 在7个基准上取得5种评估指标的领先表现
  • 适合需要高精度与实时性的视觉跟踪场景

目标跟踪是计算机视觉的基础任务,旨在跨视频帧定位感兴趣对象。扩散模型在图像生成方面表现出色,适用于解决跟踪任务中的多项需求。本文提出一种新型扩散基跟踪方法:首先,其条件生成过程可注入目标对象信息;其次,扩散机制能自然建模时序对应关系,实现真实视频帧的重建。然而,现有扩散模型依赖大量不必要的高斯噪声映射,效率低下。本文提出的插值机制借鉴经典图像处理技术,提供更可解释、更稳定且更快的解决方案,专为跟踪任务优化。DINTR 结合扩散模型优势并规避其局限,在七个基准上覆盖五种评估指标,表现卓越。

原文摘要 · Abstract (English)

Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work proposes a novel diffusion-based methodology to formulate the tracking task. Firstly, their conditional process allows for injecting indications of the target object into the generation process. Secondly, diffusion mechanics can be developed to inherently model temporal correspondences, enabling the reconstruction of actual frames in video. However, existing diffusion models rely on extensive and unnecessary mapping to a Gaussian noise domain, which can be replaced by a more efficient and stable interpolation process. Our proposed interpolation mechanism draws inspiration from classic image-processing techniques, offering a more interpretable, stable, and faster approach tailored specifically for the object tracking task. By leveraging the strengths of diffusion models while circumventing their limitations, our Diffusion-based INterpolation TrackeR (DINTR) presents a promising new paradigm and achieves a superior multiplicity on seven benchmarks across five indicator representations.

目标跟踪扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。