用格子玻尔兹曼模型学习像素动态,提升视觉跟踪实时性与适应性。
Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- 将目标像素分解为动态晶格,通过碰撞-流运机制建模运动状态。
- 在TAP-Vid、RoboTAP等基准上实现在线实时跟踪,精度显著优于现有方法。
- 适合需要快速适应真实场景的视觉跟踪任务,如机器人导航与视频分析。
本文提出格子玻尔兹曼模型(LBM),用于学习真实世界中的像素动态以实现视觉跟踪。LBM将视觉表征分解为动态像素晶格,并通过碰撞-流运过程求解像素运动状态。具体地,通过多层预测-更新网络获取目标像素的高维分布,以估计其位置与可见性。预测阶段在目标像素的空间邻域内建模晶格碰撞,并在时间上下文中实现晶格流运;更新阶段利用在线视觉表示修正像素分布。相比现有方法,LBM展现出在线与实时的应用能力,能高效适应真实世界的视觉跟踪任务。在TAP-Vid与RoboTAP等真实点跟踪基准上的全面评估验证了LBM的效率;在TAO、BFT与OVT-B等大规模开放世界目标跟踪基准上的通用评估进一步证明了其在真实场景中的实用性。
原文摘要 · Abstract (English)
This work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking. LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes. Specifically, the high-dimensional distribution of the target pixels is acquired through a multilayer predict-update network to estimate the pixel positions and visibility. The predict stage formulates lattice collisions among the spatial neighborhood of target pixels and develops lattice streaming within the temporal visual context. The update stage rectifies the pixel distributions with online visual representations. Compared with existing methods, LBM demonstrates practical applicability in an online and real-time manner, which can efficiently adapt to real-world visual tracking tasks. Comprehensive evaluations of real-world point tracking benchmarks such as TAP-Vid and RoboTAP validate LBM's efficiency. A general evaluation of large-scale open-world object tracking benchmarks such as TAO, BFT, and OVT-B further demonstrates LBM's real-world practicality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。