arXiv:2605.07642cs.CV2026-05

用多模态大模型预测第一视角手部动作,更准更鲁棒。

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

论文配图:EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
图 1 · 摘自论文原文
  • 融合视觉语言动作模型与第一视角视频编码器,联合建模运动与上下文。
  • 在EgoExo4D上达到新基准,严重视角变化下仍保持稳定性能。
  • 支持语言指令控制预测,适合人机交互与虚拟现实应用。

从第一视角视频中预测未来3D手部姿态序列对于理解人类意图及实现沉浸式应用(如AR/VR辅助与人机交互)至关重要。然而,该任务极具挑战性,因第一视角手部运动受复杂意图驱动,具有高度灵巧的关节运动,并在自我运动引发的剧烈视角变化下观测。本文提出EggHand,一个基于基础模型的第一视角手部姿态预测框架,统一多模态语义推理与动态运动建模。该方法结合来自视觉-语言-动作(VLA)模型的动作解码器,捕捉手部运动的结构化时序特性,以及从大规模第一人称视频学习到的视角感知上下文信息的视频-文本编码器。两者协同克服通用视觉编码器在自我运动下的脆弱性,实现无需人体姿态或外部追踪的运动、上下文与高层意图联合推理。在EgoExo4D数据集上的实验表明,EggHand在预测精度上达到新基准,对严重自我运动具有鲁棒性,并可通过语言任务提示实现可控预测。

原文摘要 · Abstract (English)

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly challenging problem because egocentric hand motion is driven by complex human intent, exhibits highly dexterous articulations, and is observed under drastic viewpoint shifts induced by ego-motion. In this work, we introduce EggHand, a foundation-model-based framework for egocentric hand pose forecasting that unifies multimodal semantic reasoning with dynamic motion modeling. Our approach couples an action decoder from a Vision-Language-Action (VLA) model, which captures the structured temporal dynamics of hand motion, with an egocentric video-text encoder that provides viewpoint-aware contextual information learned from large-scale first-person video. Together, these components overcome the brittleness of generic visual encoders under ego-motion and enable joint reasoning over motion, context, and high-level intent-without relying on body pose or external tracking. Experiments on the EgoExo4D dataset show that EggHand sets a new state of the art in forecasting accuracy, remains robust under severe ego-motion, and further enables controllable prediction via language-based task prompts. Project page: https://jyoun9.github.io/EggHand

手部姿态多模态第一视角生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。