arXiv:2512.02006cs.CV2025-12被引 3

多视角视频中追踪任意点,利用视角间信息提升轨迹准确性。

MV-TAP: Tracking Any Point in Multi-View Videos

  • 融合相机几何与跨视角注意力机制,聚合多视角时空信息。
  • 在挑战性基准上超越现有方法,显著提升轨迹估计可靠性。
  • 适合多摄像头系统中的动态物体追踪研究者使用。

多视角摄像头系统可对复杂现实场景进行丰富观测,理解多视角下动态物体已成为诸多应用的核心。本文提出MV-TAP,一种新型点追踪方法,通过利用跨视角信息,在动态场景的多视角视频中追踪点。该方法结合相机几何结构与跨视角注意力机制,聚合多视角的时空信息,实现更完整、更可靠的轨迹估计。为支持该任务,我们构建了大规模合成训练数据集及真实世界评估数据集。大量实验表明,MV-TAP在具有挑战性的基准测试中优于现有点追踪方法,为多视角点追踪研究建立了有效基线。

原文摘要 · Abstract (English)

Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications. In this work, we present MV-TAP, a novel point tracker that tracks points across multi-view videos of dynamic scenes by leveraging cross-view information. MV-TAP utilizes camera geometry and a cross-view attention mechanism to aggregate spatio-temporal information across views, enabling more complete and reliable trajectory estimation in multi-view videos. To support this task, we construct a large-scale synthetic training dataset and real-world evaluation sets tailored for multi-view tracking. Extensive experiments demonstrate that MV-TAP outperforms existing point-tracking methods on challenging benchmarks, establishing an effective baseline for advancing research in multi-view point tracking.

多视角追踪点追踪跨视角注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。