arXiv:2602.10659cs.CV2026-02被引 2

用多模态先验生成更自然的人物交互动作,解决动作不顺、物体不动、互动弱三大难题。

Multimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation

  • 引入图文姿态等多模态数据作先验,指导动作生成
  • 通过几何关键点和动态属性增强物体表征,提升真实感
  • 分阶段扩散模型+交互监督,强化人与物体的协同动作

针对文本驱动的3D人物-物体交互(HOI)动作生成难题,现有方法依赖直接文本到交互映射,因跨模态差距导致三大缺陷:(Q1)人体动作欠优,(Q2)物体运动不自然,(Q3)人机互动薄弱。为此,提出MP-HOI框架,基于四项核心洞察:(1)利用大模型中的文本、图像、姿态/物体多模态数据作为先验,优化数据建模;(2)融合几何关键点、接触特征与动态属性,增强物体表征;(3)设计模态感知的混合专家(MoE)模型,实现高效多模态特征融合;(4)采用分层扩散框架,通过专项监督逐步优化人机交互特征。大量实验表明,MP-HOI在生成高保真、细粒度的HOI动作上优于现有方法。

原文摘要 · Abstract (English)

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the significant cross-modality gap: (Q1) sub-optimal human motion, (Q2) unnatural object motion, and (Q3) weak interaction between humans and objects. To address these challenges, we propose MP-HOI, a novel framework grounded in four core insights: (1) Multimodal Data Priors: We leverage multimodal data (text, image, pose/object) from large multimodal models as priors to guide HOI generation, which tackles Q1 and Q2 in data modeling. (2) Enhanced Object Representation: We improve existing object representations by incorporating geometric keypoints, contact features, and dynamic properties, enabling expressive object representations, which tackles Q2 in data representation. (3) Multimodal-Aware Mixture-of-Experts (MoE) Model: We propose a modality-aware MoE model for effective multimodal feature fusion paradigm, which tackles Q1 and Q2 in feature fusion. (4) Cascaded Diffusion with Interaction Supervision: We design a cascaded diffusion framework that progressively refines human-object interaction features under dedicated supervision, which tackles Q3 in interaction refinement. Comprehensive experiments demonstrate that MP-HOI outperforms existing approaches in generating high-fidelity and fine-grained HOI motions.

3D生成人机交互扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。