让虚拟人能自然地与可动物体互动,且泛化能力强。
GIRAF: Towards Generalizable Human Interactions with Articulated Objects

- 用物体为中心的表示法统一手物接触与物体表面,实现协同建模。
- 在未见过的物体位置和形状上表现更优,超越现有方法。
- 适合机器人训练、虚拟角色生成等需要真实交互的场景。
生成逼真的全身人类与可动物体交互是具身智能与图形学的核心挑战,应用于机器人训练和虚拟代理。现有模型受限:部分仅处理静态物体上的简单动作,另一些仅关注手部操作。这使得协调式全身运动——包括接近、操控并移动可动物体——仍难以实现。核心难点在于联合推理行走、精细接触与物体关节运动。模型需捕捉跨不同物体几何形状的手物对应关系,并实现从导航到操作的平滑过渡。同时,大规模配对的动作-场景数据稀缺,制约了在多样物体位置与形状下的泛化能力。我们提出一种文本条件扩散模型,通过三项核心设计解决上述问题:以物体为中心的表示法统一手物接触与物体表面;混合域训练策略平衡行走与交互学习;基于接触的增强方案提升训练多样性。实验表明,该方法在未见物体配置下具备强泛化能力,显著优于当前最先进方法。
原文摘要 · Abstract (English)
Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, while others restrict attention to hand-only manipulation. This leaves open the problem of generating coordinated full-body motion that approaches, manipulates, and moves articulated objects in a realistic and generalizable way. The key difficulty lies in reasoning jointly about locomotion, fine-grained contact, and object articulation. Models must capture subtle hand-object correspondences that transfer across object geometries, while also producing seamless transitions from navigation to manipulation. At the same time, the scarcity of large-scale paired motion-scene data makes it difficult to generalize across diverse object positions and shapes. We introduce a text-conditioned diffusion model that addresses these challenges through three core ideas: an object-centric representation that unifies hand-object contact with object surfaces, a mixed-domain training strategy that balances locomotion and interaction, and a contact-based augmentation scheme that expands training diversity. Through experiments, our method demonstrated strong generalization to unseen object configurations, surpassing current state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。