arXiv:2605.22272cs.ROcs.CV2026-05被引 1

零样本生成人形与物体交互动作,无需3D模型

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

论文配图:Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
图 1 · 摘自论文原文
  • 用4D点轨迹统一表示机器人与物体运动,避免几何错位
  • 仅追踪关键点(底座、手部、物体),跳过复杂重定向过程
  • 结合行为基础模型隐空间,实现自然步态的零样本部署

全身人形-物体交互受限于高质量3D数据的稀缺。尽管视频生成先验提供了有前景的替代方案,但现有方法因依赖几何先验(如显式CAD模型)而存在表示错位问题,且因大量形态变换和形貌不匹配导致重定向复杂。我们提出Imagine2Real,一种无需3D几何信息的零样本人形-物体交互框架。为解决错位问题,将机器人与物体运动统一建模为4D点轨迹;为克服重定向复杂性,采用关键点追踪器仅跟踪底座、手部及物体等稀疏关键点,彻底绕过易出错的重定向流程。为在稀疏信号下保持自然步态,利用行为基础模型(BFM)的隐空间作为追踪器搜索域。通过渐进式训练策略,Imagine2Real以简单追踪奖励学习鲁棒行为,可在动作捕捉(mocap)系统中实现零样本物理部署。

原文摘要 · Abstract (English)

Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system.

人形交互视频生成零样本行为模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。