arXiv:2603.13615cs.CVcs.RO2026-03被引 5

无需未来物体状态,仅凭动作信号生成真实感手物交互

Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis

  • 从动作信号出发,构建无未来信息的物理一致第一人称交互模型
  • 在HOT3D数据集上超越强基线,实现更真实的接触一致性生成
  • 适合需要真实物理模拟的具身智能数据生成场景

为服务具身AI的可扩展数据生成,世界模型应作为真正的模拟器,仅依据用户动作推断交互动态,而非依赖已知未来物体状态的条件视频生成。在此背景下,第一人称人体-物体交互(HOI)世界模型对预测物理合理的第一人称轨迹至关重要。然而,由于快速头部运动、严重遮挡及高自由度手部动作导致接触拓扑突变,构建此类模型极具挑战。现有方法常通过访问未来物体轨迹绕开物理难题。本文提出EgoHOI,一种不依赖未来状态的首人称HOI世界模型,可仅凭动作信号生成逼真且接触一致的交互。通过将3D估计中的几何与运动先验提炼为物理感知嵌入,该模型在无未来信息条件下正则化第一人称轨迹,使其符合物理规律。在HOT3D数据集上的实验表明,EgoHOI持续优于强基线,消融实验验证了其物理感知设计的有效性。

原文摘要 · Abstract (English)

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object states. In this context, egocentric Human-Object Interaction (HOI) world models are critical for predicting physically grounded first-person rollouts. However, building such models is profoundly challenging due to rapid head motions, severe occlusions, and high-DoF hand articulations that abruptly alter contact topologies. Consequently, existing approaches often circumvent these physics challenges by resorting to conditional video generation with access to known future object trajectories. We introduce EgoHOI, an egocentric HOI world model that breaks away from this shortcut to simulate photorealistic, contact-consistent interactions from action signals alone. To ensure physical accuracy without future-state inputs, EgoHOI distills geometric and kinematic priors from 3D estimates into physics-informed embeddings. These embeddings regularize the egocentric rollouts toward physically valid dynamics. Experiments on the HOT3D dataset demonstrate consistent gains over strong baselines, and ablations validate the effectiveness of our physics-informed design.

第一人称交互物理模拟手物交互世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。