arXiv:2512.16907cs.CVcs.AI2025-12被引 8

用视觉语言推理预测手部动作轨迹,提升真实场景泛化能力

Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos

  • 构建视觉语言接口,让推理与动作生成协同优化
  • 在21.9万条6自由度轨迹上实现阶段感知的精准预测
  • 适合做交互理解与具身智能研究的学者参考

现有3D手部轨迹预测研究受限于将动作与语义监督分离的数据集,以及推理与行为关联薄弱的模型。为此,我们提出EgoMAN数据集,这是一个大规模的自视角交互数据集,包含21.9万条6自由度轨迹和300万个结构化问答对,支持语义、空间与运动推理。随后,我们设计EgoMAN模型,一种通过轨迹-令牌接口连接视觉语言推理与运动生成的框架。该模型分阶段训练,使推理与运动动态对齐,在真实场景中展现出良好泛化性,能生成准确且阶段感知的手部轨迹。

原文摘要 · Abstract (English)

Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes.

3D轨迹预测视觉语言模型具身智能自视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。