arXiv:2503.20172cs.CV2025-03CVPR被引 23

用精细几何建模提升人物交互生成真实感

Guiding Human-Object Interactions with Rich Geometry and Relations

  • 基于物体网格的关键点构建交互距离场,保留复杂几何细节
  • 引入时空注意力关系模型,显著提升动作语义一致性
  • 适合虚拟现实、动画生成等需高保真交互的场景

人-物体交互(HOI)合成对虚拟现实等应用至关重要。现有方法多依赖物体中心点或最近点等简化表示,难以捕捉几何复杂性,导致交互真实性不足。为此,本文提出ROG——一种基于扩散模型的框架,通过从物体网格中选取边界聚焦与细节关键点,实现对物体几何的全面表征,并构建交互距离场(IDF)以捕捉稳健的HOI动态。进一步设计了融合空间与时间注意力的关系模型,增强对复杂交互关系的理解,从而优化生成动作的IDF,指导运动生成过程,实现关系感知且语义一致的动作输出。实验表明,ROG在合成HOI的真实性和语义准确性上显著优于现有最优方法。

原文摘要 · Abstract (English)

Human-object interaction (HOI) synthesis is crucial for creating immersive and realistic experiences for applications such as virtual reality. Existing methods often rely on simplified object representations, such as the object's centroid or the nearest point to a human, to achieve physically plausible motions. However, these approaches may overlook geometric complexity, resulting in suboptimal interaction fidelity. To address this limitation, we introduce ROG, a novel diffusion-based framework that models the spatiotemporal relationships inherent in HOIs with rich geometric detail. For efficient object representation, we select boundary-focused and fine-detail key points from the object mesh, ensuring a comprehensive depiction of the object's geometry. This representation is used to construct an interactive distance field (IDF), capturing the robust HOI dynamics. Furthermore, we develop a diffusion-based relation model that integrates spatial and temporal attention mechanisms, enabling a better understanding of intricate HOI relationships. This relation model refines the generated motion's IDF, guiding the motion generation process to produce relation-aware and semantically aligned movements. Experimental evaluations demonstrate that ROG significantly outperforms state-of-the-art methods in the realism and semantic accuracy of synthesized HOIs.

人机交互扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。