arXiv:2601.07154cs.CV2026-01

从第一视角视频中实时识别运动意图,助力体育与快速动作场景分析。

Motion Focus Recognition in Fast-Moving Egocentric Video

  • 基于相机位姿基础模型,结合滑动批处理实现高效推理。
  • 在自建数据集上实现实时性能,内存占用可控。
  • 适合边缘部署,为快速运动场景提供新分析视角。

从视觉-语言-动作系统到机器人应用,现有第一视角数据集主要关注动作识别任务,而忽视了运动分析在体育及其他快速运动场景中的关键作用。为填补这一空白,我们提出一种实时运动焦点识别方法,可从任意第一视角视频中估计主体的移动意图。该方法利用相机位姿估计的基础模型,并引入系统级优化,实现高效且可扩展的推理。在自建的第一视角动作数据集上,通过滑动批处理推理策略,本方法实现了实时性能并保持了可管理的内存消耗。该工作使以运动为中心的分析适用于边缘部署,并为体育和快速运动活动研究提供了互补视角。

原文摘要 · Abstract (English)

From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the subject's locomotion intention from any egocentric video. We leverage the foundation model for camera pose estimation and introduce system-level optimizations to enable efficient and scalable inference. Evaluated on a collected egocentric action dataset, our method achieves real-time performance with manageable memory consumption through a sliding batch inference strategy. This work makes motion-centric analysis practical for edge deployment and offers a complementary perspective to existing egocentric studies on sports and fast-movement activities.

第一视角运动识别实时推理边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。