arXiv:2506.00827cs.CV2025-06被引 1

通过手部聚焦稳定视频,提升第一人称动作识别准确率

Improving Keystep Recognition in Ego-Video via Dexterous Focus

  • 将第一人称视频转为手部聚焦的稳定画面
  • 在Ego-Exo4D数据集上超越现有基线方法
  • 无需改动模型结构,通用性强

本文针对第一人称视角下人类活动理解的挑战。由于许多动作中头部高度动态,传统活动识别技术面临困难。我们提出一种独立于网络架构的框架,通过将第一人称视频输入限制为稳定且聚焦手部的视频来应对这些挑战。实验表明,仅通过这一简单的视频转换,即可在Ego-Exo4D细粒度按键步骤识别基准上超越现有第一人称视频基线方法,且无需修改底层模型结构。

原文摘要 · Abstract (English)

In this paper, we address the challenge of understanding human activities from an egocentric perspective. Traditional activity recognition techniques face unique challenges in egocentric videos due to the highly dynamic nature of the head during many activities. We propose a framework that seeks to address these challenges in a way that is independent of network architecture by restricting the ego-video input to a stabilized, hand-focused video. We demonstrate that this straightforward video transformation alone outperforms existing egocentric video baselines on the Ego-Exo4D Fine-Grained Keystep Recognition benchmark without requiring any alteration of the underlying model infrastructure.

第一人称视频动作识别手部聚焦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。