arXiv:2512.19049cs.CV2025-12被引 1

分离路径规划与动作生成,实现更自然的人物-物体交互合成

Decoupled Generative Modeling for Human-Object Interaction Synthesis

  • 先生成无预设路径点的轨迹,再基于轨迹合成精细动作
  • 在两个基准上优于现有方法,减少动作不同步和穿模问题
  • 适合需要真实交互动画的影视、机器人控制场景

人-物体交互(HOI)合成对3D计算机视觉和机器人技术至关重要,支撑动画与具身控制。现有方法常需人工指定中间路径点,并将所有优化目标集中于单一网络,导致复杂度高、灵活性差,易出现动作不同步或物体穿模。为此,我们提出解耦生成建模框架DecHOI,将路径规划与动作合成分离:轨迹生成器无需预设路径点生成人与物体轨迹;动作生成器基于这些路径生成详细运动。为提升接触真实感,采用聚焦远端关节动态的对抗训练策略。框架还建模运动中的交互对象,支持动态场景中响应式、长序列规划,并保持计划一致性。在FullBodyManipulation与3D-FUTURE两个基准上,DecHOI在多数定量指标与定性评估中超越现有方法,感知实验亦偏好其结果。

原文摘要 · Abstract (English)

Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control. Existing approaches often require manually specified intermediate waypoints and place all optimization objectives on a single network, which increases complexity, reduces flexibility, and leads to errors such as unsynchronized human and object motion or penetration. To address these issues, we propose Decoupled Generative Modeling for Human-Object Interaction Synthesis (DecHOI), which separates path planning and action synthesis. A trajectory generator first produces human and object trajectories without prescribed waypoints, and an action generator conditions on these paths to synthesize detailed motions. To further improve contact realism, we employ adversarial training with a discriminator that focuses on the dynamics of distal joints. The framework also models a moving counterpart and supports responsive, long-sequence planning in dynamic scenes, while preserving plan consistency. Across two benchmarks, FullBodyManipulation and 3D-FUTURE, DecHOI surpasses prior methods on most quantitative metrics and qualitative evaluations, and perceptual studies likewise prefer our results.

交互生成动作合成解耦建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。