arXiv:2506.09984cs.CVcs.AI2025-06被引 19

让多人互动视频生成更精准,通过布局对齐音频实现角色分区域控制。

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

  • 基于掩码预测自动推断多概念布局,实现条件与身份的区域绑定。
  • 在两人至三人对话视频生成中达到高质量效果,支持多参考图像定制。
  • 适合需要精细角色控制的影视动画、虚拟人交互等应用。

近年来,基于文本、图像和音频等多模态条件的端到端人体动画取得了显著进展。然而,现有方法大多仅能处理单一主体,并以全局方式注入条件,忽略了同一视频中存在多个概念且涉及丰富的人-人、人-物互动场景。这种全局假设限制了对多主体(包括人与物体)的精确身份控制,制约了实际应用。本文摒弃单主体假设,提出一种新框架,强制多模态条件与各身份时空轨迹的强区域特异性绑定。给定多个概念的参考图像,方法通过掩码预测器匹配去噪视频与各参考外观之间的视觉线索,自动推断布局信息;同时将局部音频条件注入对应区域,以迭代方式实现布局对齐的模态匹配。该设计可生成高质量的双至三人群体对话视频,或基于多参考图像进行视频定制。实验结果与消融研究验证了显式布局控制相比隐式方法及其他现有方法的有效性。视频演示见 https://zhenzhiwang.github.io/interacthuman/

原文摘要 · Abstract (English)

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios where multiple concepts could appear in the same video with rich human-human interactions and human-object interactions. Such a global assumption prevents precise and per-identity control of multiple concepts including humans and objects, therefore hinders applications. In this work, we discard the single-entity assumption and introduce a novel framework that enforces strong, region-specific binding of conditions from modalities to each identity's spatiotemporal footprint. Given reference images of multiple concepts, our method could automatically infer layout information by leveraging a mask predictor to match appearance cues between the denoised video and each reference appearance. Furthermore, we inject local audio condition into its corresponding region to ensure layout-aligned modality matching in an iterative manner. This design enables the high-quality generation of human dialogue videos between two to three people or video customization from multiple reference images. Empirical results and ablation studies validate the effectiveness of our explicit layout control for multi-modal conditions compared to implicit counterparts and other existing methods. Video demos are available at https://zhenzhiwang.github.io/interacthuman/

视频生成多主体控制音频对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。