arXiv:2412.04592cs.CV2024-12中稿 · WACV 2025被引 14

构建首个面向第一视角视频的点追踪基准,提升复杂场景下追踪精度。

EgoPoints: Advancing Point Tracking for Egocentric Videos

  • 设计新基准EgoPoints,包含4700个高难度追踪点。
  • 在遮挡与重识别场景中,模型准确率提升2.7至2.8个百分点。
  • 适合研究第一视角视觉、点追踪与自监督学习的学者参考。

我们提出EgoPoints,一个面向第一视角视频的点追踪基准。通过标注4700个具有挑战性的追踪轨迹,相比TAP-Vid-DAVIS基准,本数据集包含9倍更多的离屏点和59倍更多的需重识别点。为此,我们引入专门评估在视、离屏及重识别条件下追踪性能的指标。同时,我们提出一种生成半真实序列的流水线,利用动态Kubric物体与EPIC Fields场景点自动生成11000条序列,并提供自动标注的真值。在这些合成序列上微调点追踪模型后,在真实标注的EgoPoints数据集上评估,CoTracker在所有指标上均取得提升,其中平均追踪准确率$δ^ ext{⋆}_{ ext{avg}}$提高2.7个百分点,重识别准确率(ReID$δ_{ ext{avg}}$)提升2.4个百分点;PIP's++的对应指标分别提升0.3和2.8个百分点。

原文摘要 · Abstract (English)

We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS evaluation benchmark, we include 9x more points that go out-of-view and 59x more points that require re-identification (ReID) after returning to view. To measure the performance of models on these challenging points, we introduce evaluation metrics that specifically monitor tracking performance on points in-view, out-of-view, and points that require re-identification. We then propose a pipeline to create semi-real sequences, with automatic ground truth. We generate 11K such sequences by combining dynamic Kubric objects with scene points from EPIC Fields. When fine-tuning point tracking methods on these sequences and evaluating on our annotated EgoPoints sequences, we improve CoTracker across all metrics, including the tracking accuracy $δ^\star_{\text{avg}}$ by 2.7 percentage points and accuracy on ReID sequences (ReID$δ_{\text{avg}}$) by 2.4 points. We also improve $δ^\star_{\text{avg}}$ and ReID$δ_{\text{avg}}$ of PIPs++ by 0.3 and 2.8 respectively.

点追踪第一视角视频理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。