arXiv:2506.08797cs.CV2025-06被引 14

弱监督下实现通用人物-物体交互视频生成,支持文本控制与互动操作

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

  • 通过稀疏运动引导与双输入空间融合,解耦外观与动作信号
  • 在弱监督下实现自然交互与跨场景泛化,性能达当前最优
  • 支持文本驱动生成与交互式物体操控,适合内容创作与动画设计

为解决人物-物体交互(HOI)视频生成中的关键挑战——依赖精心标注的运动数据、对新物体/场景泛化能力有限、可及性差——我们提出HunyuanVideo-HOMA,一种弱监督多模态驱动框架。该框架通过稀疏解耦的运动引导增强可控性,降低对精确输入的依赖;将外观与运动信号编码至多模态扩散变换器(MMDiT)的双重输入空间,在共享上下文空间中融合,生成时序一致且物理合理的交互视频。为优化训练,引入基于预训练MMDiT权重初始化的参数空间HOI适配器,保留先验知识并实现高效适应;同时采用面部交叉注意力适配器,实现音视频驱动下的解剖学准确唇部同步。大量实验表明,该方法在弱监督下达到最先进的交互自然度与泛化能力。最终,HunyuanVideo-HOMA展现出在文本条件生成与交互式物体操控中的多功能性,并配备用户友好的演示界面。项目页面见 https://bone-11.github.io/homa-page/

原文摘要 · Abstract (English)

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce HunyuanVideo-HOMA, a weakly conditioned multimodal-driven framework. HunyuanVideo-HOMA enhances controllability and reduces dependency on precise inputs through sparse, decoupled motion guidance. It encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to synthesize temporally consistent and physically plausible interactions. To optimize training, we integrate a parameter-space HOI adapter initialized from pretrained MMDiT weights, preserving prior knowledge while enabling efficient adaptation, and a facial cross-attention adapter for anatomically accurate audio-driven lip synchronization. Extensive experiments confirm state-of-the-art performance in interaction naturalness and generalization under weak supervision. Finally, HunyuanVideo-HOMA demonstrates versatility in text-conditioned generation and interactive object manipulation, supported by a user-friendly demo interface. The project page is at https://bone-11.github.io/homa-page/.

视频生成多模态人机交互扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。