多模态驱动的视频生成框架,实现主体一致性与灵活条件控制。
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
- 融合图像、音频、视频和文本输入,通过跨模态融合增强理解。
- 在单/多主体场景中,身份一致性与视频真实感显著优于现有方法。
- 适合需要精准主体控制的视频生成应用,如个性化内容创作。
定制化视频生成旨在根据用户定义条件生成特定主体的视频,但现有方法常面临身份不一致和输入模态受限的问题。本文提出HunyuanCustom,一种强调主体一致性的多模态定制视频生成框架,支持图像、音频、视频和文本条件输入。基于HunyuanVideo,模型首先通过基于LLaVA的文本-图像融合模块提升多模态理解能力,并引入基于时序拼接的图像ID增强模块,强化帧间身份特征。为支持音频与视频条件生成,进一步提出专用条件注入机制:AudioNet模块通过空间交叉注意力实现层级对齐,视频驱动模块则通过基于patchify的特征对齐网络整合压缩后的条件视频特征。在单主体与多主体场景下的大量实验表明,HunyuanCustom在身份一致性、真实感及文本-视频对齐方面显著优于当前最先进的开源与闭源方法。此外,其在下游任务(如音频与视频驱动生成)中展现出良好鲁棒性。结果表明,多模态条件与身份保持策略能有效推动可控视频生成的发展。代码与模型已公开于https://hunyuancustom.github.io。
原文摘要 · Abstract (English)
Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose HunyuanCustom, a multi-modal customized video generation framework that emphasizes subject consistency while supporting image, audio, video, and text conditions. Built upon HunyuanVideo, our model first addresses the image-text conditioned generation task by introducing a text-image fusion module based on LLaVA for enhanced multi-modal understanding, along with an image ID enhancement module that leverages temporal concatenation to reinforce identity features across frames. To enable audio- and video-conditioned generation, we further propose modality-specific condition injection mechanisms: an AudioNet module that achieves hierarchical alignment via spatial cross-attention, and a video-driven injection module that integrates latent-compressed conditional video through a patchify-based feature-alignment network. Extensive experiments on single- and multi-subject scenarios demonstrate that HunyuanCustom significantly outperforms state-of-the-art open- and closed-source methods in terms of ID consistency, realism, and text-video alignment. Moreover, we validate its robustness across downstream tasks, including audio and video-driven customized video generation. Our results highlight the effectiveness of multi-modal conditioning and identity-preserving strategies in advancing controllable video generation. All the code and models are available at https://hunyuancustom.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。