用大模型指导多主体视频生成,提升一致性与个性化
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
- 利用多模态大模型理解主体间关系,避免图像与文本的模糊对应
- 在多个主体下保持时空一致,生成视频质量显著优于现有方法
- 支持可变数量主体,适合互动媒体与定制化视频创作
视频生成在深度生成模型(尤其是扩散模型)推动下取得显著进展。尽管现有方法在文本或单图生成高质量视频方面表现优异,但个性化多主体视频生成仍是未充分探索的挑战。该任务需合成包含多个独立主体的视频,每个主体由独立参考图像定义,并保证时间与空间的一致性。当前方法主要依赖将主体图像映射到文本提示中的关键词,引入歧义并限制对主体关系的有效建模。本文提出CINEMA框架,通过多模态大语言模型(MLLM)实现连贯的多主体视频生成。该方法无需显式建立主体图像与文本实体的对应关系,降低歧义性并减少标注成本。借助MLLM理解主体间关系,提升了模型可扩展性,支持使用大规模多样化数据集训练。此外,框架可适应不同数量的主体,增强个性化内容创作灵活性。大量实验表明,该方法显著提升主体一致性与整体视频连贯性,为叙事、交互媒体和个性化视频生成开辟新路径。
原文摘要 · Abstract (English)
Video generation has witnessed remarkable progress with the advent of deep generative models, particularly diffusion models. While existing methods excel in generating high-quality videos from text prompts or single images, personalized multi-subject video generation remains a largely unexplored challenge. This task involves synthesizing videos that incorporate multiple distinct subjects, each defined by separate reference images, while ensuring temporal and spatial consistency. Current approaches primarily rely on mapping subject images to keywords in text prompts, which introduces ambiguity and limits their ability to model subject relationships effectively. In this paper, we propose CINEMA, a novel framework for coherent multi-subject video generation by leveraging Multimodal Large Language Model (MLLM). Our approach eliminates the need for explicit correspondences between subject images and text entities, mitigating ambiguity and reducing annotation effort. By leveraging MLLM to interpret subject relationships, our method facilitates scalability, enabling the use of large and diverse datasets for training. Furthermore, our framework can be conditioned on varying numbers of subjects, offering greater flexibility in personalized content creation. Through extensive evaluations, we demonstrate that our approach significantly improves subject consistency, and overall video coherence, paving the way for advanced applications in storytelling, interactive media, and personalized video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。