arXiv:2605.05756cs.ROcs.CV2026-05

让人体与物体交互更自然且精准接触,解决深度模型丢失几何细节的问题。

MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation

论文配图:MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
图 1 · 摘自论文原文
  • 分层设计:宏观运动流畅性与微观空间精度并重
  • 接触精确度提升,长时复杂轨迹生成能力突破
  • 适合虚拟内容生成、具身智能等需要高保真交互的场景

生成真实的3D人体-物体交互(HOI)是具身智能到虚拟内容创作的关键任务,需协调高层语义意图与底层物理约束。现有方法在语义对齐上表现良好,但难以保持精确的物体接触。我们发现一个关键现象: extit{几何遗忘}——随着扩散模型深度增加,语义特征逐渐掩盖物体几何特征,导致模型丧失对物体几何的感知。为此,我们提出MaMi-HOI,一种分层框架,融合 extbf{Ma}宏观运动流畅性与 extbf{Mi}微观空间精度。首先,提出几何感知邻近适配器(GAPA),显式重注入密集物体细节,实现残差贴合修正以保证精确接触。然而,这种强局部约束可能破坏全局动态,导致动作僵硬。为此,引入运动和谐适配器(KHA),主动将全身姿态与空间目标对齐,确保骨骼主动适应约束而不牺牲自然性。大量实验验证,MaMi-HOI同时实现自然运动与精确接触。关键在于,其扩展了长期复杂轨迹的生成能力,有效弥合3D场景中全局导航与高保真操作之间的差距。代码已开源。

原文摘要 · Abstract (English)

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical constraints. Existing methods excel at semantic alignment, however, they struggle to maintain precise object contact. We reveal a key finding termed \textit{Geometric Forgetting}: as diffusion model depth increases, semantic feature tend to overshadow object geometry feature, causing the model to lose its perception to object geometry. To address this, we propose MaMi-HOI, a hierarchical framework reconciling \textbf{Ma}cro-level kinematic fluidity with \textbf{Mi}cro-level spatial precision. First, to counteract geometric forgetting, we introduce the Geometry-Aware Proximity Adapter (GAPA), which explicitly re-injects dense object details to perform residual snapping corrections for precise contact. Nevertheless, such aggressive local enforcement can disrupt global dynamics, leading to robotic stiffness. In response, we introduce the Kinematic Harmony Adapter (KHA), which proactively aligns whole-body posture with spatial objectives, ensuring the skeleton actively accommodates constraints without compromising naturalness. Extensive experiments validate that MaMi-HOI simultaneously achieves natural motion and precise contact. Crucially, it extends generation capabilities to long-term tasks with complex trajectories, effectively bridging the gap between global navigation and high-fidelity manipulation in 3D scenes. Code is available at https://github.com/DON738110198/MaMi-HOI.git

3D交互扩散模型人体生成几何控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。