让多人视频生成更精准,身份一致且互动自然。
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
- 用视觉-文本融合模块精准对齐人物图像与文字描述。
- 3D-RoPE增强模块提升身份保持和主体间交互效果。
- 适合需要多角色一致生成的视频创作场景。
尽管视频生成技术取得进展,现有模型在多主体定制化生成中仍缺乏细粒度控制能力,难以保证身份一致性与交互合理性。本文提出PolyVivid框架,实现灵活且身份一致的多主体视频生成。为建立主体图像与文本实体间的准确对应,设计基于视觉-语言大模型(VLLM)的图文融合模块,将视觉身份嵌入文本空间以实现精确定位。为进一步强化身份保持与主体间交互,提出基于3D-RoPE的增强模块,实现文本与图像嵌入的结构化双向融合。同时开发注意力继承的身份注入模块,有效将融合后的身份特征注入视频生成过程,缓解身份漂移问题。最后构建基于多模态大模型(MLLM)的数据流水线,结合MLLM驱动的定位、分割及团状主体整合策略,生成高质量多主体数据,显著提升主体区分度并降低下游生成中的歧义性。大量实验表明,PolyVivid在身份保真度、视频真实感与主体对齐性方面均优于现有开源及商业基线模型。
原文摘要 · Abstract (English)
Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and identity-consistent generation. To establish accurate correspondences between subject images and textual entities, we design a VLLM-based text-image fusion module that embeds visual identities into the textual space for precise grounding. To further enhance identity preservation and subject interaction, we propose a 3D-RoPE-based enhancement module that enables structured bidirectional fusion between text and image embeddings. Moreover, we develop an attention-inherited identity injection module to effectively inject fused identity features into the video generation process, mitigating identity drift. Finally, we construct an MLLM-based data pipeline that combines MLLM-based grounding, segmentation, and a clique-based subject consolidation strategy to produce high-quality multi-subject data, effectively enhancing subject distinction and reducing ambiguity in downstream video generation. Extensive experiments demonstrate that PolyVivid achieves superior performance in identity fidelity, video realism, and subject alignment, outperforming existing open-source and commercial baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。