arXiv:2410.20395cs.CVeess.IV2024-10中稿 · ance at the Asian …被引 2

用单目深度信息提升视觉追踪鲁棒性,尤其应对运动模糊和目标丢失。

Depth Attention for Robust RGB Tracking

  • 引入深度注意力机制,无需RGB-D相机即可融合深度信息。
  • 在六个基准上实现新最优性能,显著提升模糊与遮挡场景下的追踪准确率。
  • 适合关注真实场景下视觉追踪鲁棒性的研究者与工程师。

RGB视频目标追踪是计算机视觉的基础任务。利用深度信息可有效提升追踪效果,尤其在处理运动模糊目标时。然而,常用追踪基准中常缺少深度信息。本文提出一种新框架,通过单目深度估计来应对目标丢失或受运动模糊影响的问题。首次提出深度注意力机制,并设计简单框架,实现深度信息与主流追踪算法的无缝集成,无需RGB-D相机。在六个挑战性基准上进行充分实验,结果表明本方法在多个强基线基础上持续提升性能,达到新的最佳水平。我们相信该方法将为真实场景下更复杂的视觉追踪方案开辟新可能。代码与模型已公开:https://github.com/LiuYuML/Depth-Attention。

原文摘要 · Abstract (English)

RGB video object tracking is a fundamental task in computer vision. Its effectiveness can be improved using depth information, particularly for handling motion-blurred target. However, depth information is often missing in commonly used tracking benchmarks. In this work, we propose a new framework that leverages monocular depth estimation to counter the challenges of tracking targets that are out of view or affected by motion blur in RGB video sequences. Specifically, our work introduces following contributions. To the best of our knowledge, we are the first to propose a depth attention mechanism and to formulate a simple framework that allows seamlessly integration of depth information with state of the art tracking algorithms, without RGB-D cameras, elevating accuracy and robustness. We provide extensive experiments on six challenging tracking benchmarks. Our results demonstrate that our approach provides consistent gains over several strong baselines and achieves new SOTA performance. We believe that our method will open up new possibilities for more sophisticated VOT solutions in real-world scenarios. Our code and models are publicly released: https://github.com/LiuYuML/Depth-Attention.

目标追踪深度感知注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。