arXiv:2608.03158cs.CV2026-08

用表面关键点轨迹表示物体运动,实现多物与关节动作的协同生成。

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

论文配图:Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
图 1 · 摘自论文原文
  • 以表面关键点轨迹替代传统关节建模,直接从点运动捕捉物体动态。
  • 在多个数据集上生成效果优于或相当现有方法,涵盖单物、多物及关节物体场景。
  • 适合需要复杂人物交互生成的场景,如动画、虚拟现实和机器人仿真。

日常活动需要人类全身运动与周围物体运动协调配合。尽管人-物体交互(HOI)生成已有进展,但多数方法仅适用于单一刚性物体,难以扩展到多物体或具有多样关节机制的可动物体场景。本文提出表面关键点轨迹作为物体运动表示:对每个刚性部件(无论是独立物体还是可动结构的一部分),追踪一组非共线的表面点随时间变化。该表示无需显式指定关节类型,即可直接处理多物体协调与多种关节机制。为建模身体各部位与物体接触的时间与空间位置,引入时空接触距离场,将基于距离的接触建模扩展至全身、多物体及可动结构场景。我们将HOI生成分解为三个阶段:从文本或路径点生成物体运动,预测接触距离场,最后通过接触引导优化合成全身动作。在ParaHome、HIMO、ARCTIC和OMOMO数据集上的实验表明,本方法在单物体、多物体及可动物体交互设置下均达到或优于现有方法的表现。

原文摘要 · Abstract (English)

Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.

人-物交互动作生成可动物体表面点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。