arXiv:2507.16015cs.CV2025-07ICCV被引 2

拆解第一人称视觉跟踪难在哪,发现更多是任务本身而非视角导致。

Is Tracking really more challenging in First Person Egocentric Vision?

  • 设计新评测框架,分离视角与人类动作场景的影响
  • 发现多数挑战源于人机交互任务本身,非第一人称视角独有
  • 为后续研究指明应聚焦具体任务难点,而非盲目优化视角适应性

在第一人称视角(egocentric vision)中,视觉目标跟踪与分割已成为理解人类活动的基础任务。近期研究通过对比不同场景的基准测试,认为第一人称视角相较于以往研究领域更具挑战性。然而,这些结论基于显著不同的实验场景,而许多被归因于第一人称视角的困难特征,在第三人称的人机交互视频中同样存在。这引发了一个关键问题:性能下降究竟源于第一人称视角的特殊性,还是更广泛的‘人机交互’任务本身的复杂性?为此,我们提出一项新的基准研究,旨在解耦上述因素。我们的评估策略能够更精确地区分第一人称视角相关挑战与人机交互任务固有挑战。通过这一分析,我们揭示了第一人称跟踪与分割真正困难的根源,为该任务的针对性改进提供了更清晰的方向。

原文摘要 · Abstract (English)

Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person egocentric vision presents challenges compared to previously studied domains. However, these claims are based on evaluations conducted across significantly different scenarios. Many of the challenging characteristics attributed to egocentric vision are also present in third person videos of human-object activities. This raises a critical question: how much of the observed performance drop stems from the unique first person viewpoint inherent to egocentric vision versus the domain of human-object activities? To address this question, we introduce a new benchmark study designed to disentangle such factors. Our evaluation strategy enables a more precise separation of challenges related to the first person perspective from those linked to the broader domain of human-object activity understanding. By doing so, we provide deeper insights into the true sources of difficulty in egocentric tracking and segmentation, facilitating more targeted advancements on this task.

目标跟踪第一人称视觉评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。