arXiv:2508.19852cs.CV2025-08被引 4

基于手部轨迹预测动作与视觉未来,实现人机交互的统一建模。

Ego-centric Predictive Model Conditioned on Hand Trajectories

  • 分两阶段建模:先预测手部轨迹,再用因果交叉注意力引导扩散模型生成视频。
  • 在Ego4D等数据集上,动作预测和视频生成均优于现有方法。
  • 适合机器人规划与人类行为理解场景,可同时输出动作和视觉结果。

在第一人称场景中,同时预测下一步动作及其视觉结果对理解人机交互和实现机器人规划至关重要。然而,现有方法难以联合建模这两方面:视觉-语言-动作(VLA)模型关注动作预测,但未显式建模动作对视觉场景的影响;视频预测模型虽能生成未来帧,却未基于特定动作条件,常导致不现实或上下文不一致的结果。为此,我们提出一个统一的两阶段预测框架,以手部轨迹为条件,联合建模动作与视觉未来。第一阶段通过连续状态建模处理异构输入(视觉观测、语言、动作历史),显式预测未来手部轨迹。第二阶段引入因果交叉注意力融合多模态线索,利用推断出的动作信号引导基于潜在空间的扩散模型(LDM),逐帧生成未来视频。该方法是首个专为第一人称人类活动理解与机器人操作任务设计的统一模型,明确预测未来动作及其视觉后果。在Ego4D、BridgeData和RLBench上的大量实验表明,本方法在动作预测与未来视频合成上均超越现有最先进基线。

原文摘要 · Abstract (English)

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these aspects. Vision-Language-Action (VLA) models focus on action prediction but lack explicit modeling of how actions influence the visual scene, while video prediction models generate future frames without conditioning on specific actions, often resulting in implausible or contextually inconsistent outcomes. To bridge this gap, we propose a unified two-stage predictive framework that jointly models action and visual future in egocentric scenarios, conditioned on hand trajectories. In the first stage, we perform consecutive state modeling to process heterogeneous inputs (visual observations, language, and action history) and explicitly predict future hand trajectories. In the second stage, we introduce causal cross-attention to fuse multi-modal cues, leveraging inferred action signals to guide an image-based Latent Diffusion Model (LDM) for frame-by-frame future video generation. Our approach is the first unified model designed to handle both egocentric human activity understanding and robotic manipulation tasks, providing explicit predictions of both upcoming actions and their visual consequences. Extensive experiments on Ego4D, BridgeData, and RLBench demonstrate that our method outperforms state-of-the-art baselines in both action prediction and future video synthesis.

动作预测视频生成扩散模型第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。