通过跨模态推理实现视频生成中主体一致性,提升复杂场景下的视觉连贯性。
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- 用多模态大模型解析提示词中的主体关系与动态交互,生成主体感知的隐状态。
- 在OpenS2V基准上,主体一致性、自然度和文本相关性均优于现有开源与商用模型。
- 适用于单主体到多主体异构实体的复杂视频生成场景,适合影视创作与虚拟仿真。
扩散变压器在生成高保真视频方面表现出色,能够呈现视觉连贯的帧和丰富的细节。然而,现有视频生成模型在主体一致性方面仍存在不足,主要源于难以解析包含复杂空间关系、时间逻辑和多个主体间交互的提示。为此,我们提出BindWeave,一个统一框架,可处理从单主体到包含异构实体的复杂多主体场景的广泛视频生成任务。为将复杂提示语义绑定至具体视觉主体,我们引入了MLLM-DiT框架:预训练的多模态大语言模型进行深度跨模态推理,实现实体定位、角色与属性解耦及交互识别,生成主体感知的隐状态,用于引导扩散变压器实现高保真、主体一致的视频生成。在OpenS2V基准上的实验表明,该方法在主体一致性、自然度和文本相关性方面均显著优于现有开源与商业模型。
原文摘要 · Abstract (English)
Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in parsing prompts that specify complex spatial relationships, temporal logic, and interactions among multiple subjects. To address this issue, we propose BindWeave, a unified framework that handles a broad range of subject-to-video scenarios from single-subject cases to complex multi-subject scenes with heterogeneous entities. To bind complex prompt semantics to concrete visual subjects, we introduce an MLLM-DiT framework in which a pretrained multimodal large language model performs deep cross-modal reasoning to ground entities and disentangle roles, attributes, and interactions, yielding subject-aware hidden states that condition the diffusion transformer for high-fidelity subject-consistent video generation. Experiments on the OpenS2V benchmark demonstrate that our method achieves superior performance across subject consistency, naturalness, and text relevance in generated videos, outperforming existing open-source and commercial models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。