HuMo统一生成人像视频,支持文图音三模态协同控制。
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
- 构建了图文音配对数据集,解决多模态训练数据稀缺问题。
- 分阶段训练实现主体保持与音画同步,性能超越专用方法。
- 支持灵活可控的时序自适应引导,适合高精度视频生成任务。
人像视频生成(HCVG)旨在从文本、图像和音频等多模态输入中合成人物视频。现有方法因两个挑战难以有效协调异构模态:一是缺乏配对的三元组训练数据,二是难以协同主体保持与音画同步任务。本文提出统一框架HuMo,针对第一个挑战,构建高质量、多样化的图文音配对数据集;针对第二个挑战,设计两阶段渐进式多模态训练范式。在主体保持任务中,采用最小侵入式图像注入策略以保留基础模型的提示遵循与视觉生成能力;在音画同步任务中,除常规音频交叉注意力外,提出“先预测后聚焦”策略,隐式引导模型将音频与面部区域关联。为联合学习多模态可控性,基于已有能力逐步引入音画同步任务。推理时,设计时间自适应的无分类器引导策略,动态调整去噪步骤中的引导权重。大量实验证明,HuMo在子任务上优于专用最先进方法,建立了统一的多模态协同控制框架。
原文摘要 · Abstract (English)
Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of training data with paired triplet conditions and the difficulty of collaborating the sub-tasks of subject preservation and audio-visual sync with multimodal inputs. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct a high-quality dataset with diverse and paired text, reference images, and audio. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies. For the subject preservation task, to maintain the prompt following and visual generation abilities of the foundation model, we adopt the minimal-invasive image injection strategy. For the audio-visual sync task, besides the commonly adopted audio cross-attention layer, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multimodal inputs, building on previously acquired capabilities, we progressively incorporate the audio-visual sync task. During inference, for flexible and fine-grained multimodal control, we design a time-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG. Project Page: https://phantom-video.github.io/HuMo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。