arXiv:2607.01768cs.CV2026-07被引 1

联合生成手物接触图,让动作更真实自然。

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation

论文配图:JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
图 1 · 摘自论文原文
  • 单阶段扩散模型同时生成手物动作与动态接触图。
  • 在GRAB和ARCTIC数据集上显著减少穿透和漂浮问题。
  • 适合需要逼真交互的虚拟现实与机器人应用。

文本驱动的手物交互(HOI)生成在沉浸式应用与机器人领域受到关注,但生成物理合理的交互仍具挑战。即使动作看似自然,微小的接触误差也会导致悬浮、穿模等明显瑕疵。现有方法依赖显式接触提示或隐式抓握先验,通常采用多阶段流水线,无法建模随时间演化的接触关系。我们提出JointHOI,一种单阶段扩散框架,从文本直接联合生成3D手物运动与基于距离的动态接触图。通过将接触视为辅助内部模态,联合生成使模型在训练中学习接触与运动的耦合关系。推理时,接触引导采样确保生成的接触图与运动所暗示的几何一致性,提升时间稳定性,降低穿模与悬浮现象。在GRAB和ARCTIC数据集上的实验表明,相比先前方法,本方法在文本契合度与物理合理性方面均有持续提升。

原文摘要 · Abstract (English)

Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit grasp priors, but typically rely on multi stage pipelines and fail to model temporally evolving contact. We present JointHOI, a single stage diffusion framework that jointly generates 3D hand object motion and dynamic, distance based contact maps from text. By treating contact as an auxiliary inner modality, joint generation enables the model to learn contact motion coupling during training. At inference, contact guided sampling enforces consistency between generated contact maps and motion implied geometry, improving temporal stability and reducing penetration and floating. Experiments on GRAB and ARCTIC demonstrate consistent improvements in text adherence and physical plausibility over prior methods.

手物交互扩散模型3D生成物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。