arXiv:2504.19894cs.CV2025-04被引 6

用分阶段方法生成连贯电影场景,支持多角色与复杂交互。

CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition

  • 先用大模型生成场景计划,再微调图像模型合成关键帧。
  • 在自建数据集上实现画面连贯、内容丰富的电影级输出。
  • 适合影视生成、创意设计及视频内容创作研究者使用。

我们提出CineVerse,一种面向电影场景构图的新框架。与传统多镜头生成类似,该任务强调帧间的一致性与连续性,同时聚焦于电影制作中的挑战,如多角色、复杂互动和视觉特效。为学习生成此类内容,我们首先构建了CineVerse数据集,并基于此训练两阶段方法:第一阶段,利用大语言模型(LLM)根据高层场景描述生成详细的环境设定、角色配置及分镜计划;第二阶段,微调文本到图像生成模型以合成高质量的视觉关键帧。实验表明,CineVerse在生成视觉连贯且语境丰富的电影场景方面取得显著提升,为电影视频合成的进一步探索奠定了基础。

原文摘要 · Abstract (English)

We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on addressing challenges inherent to filmmaking, such as multiple characters, complex interactions, and visual cinematic effects. In order to learn to generate such content, we first create the CineVerse dataset. We use this dataset to train our proposed two-stage approach. First, we prompt a large language model (LLM) with task-specific instructions to take in a high-level scene description and generate a detailed plan for the overall setting and characters, as well as the individual shots. Then, we fine-tune a text-to-image generation model to synthesize high-quality visual keyframes. Experimental results demonstrate that CineVerse yields promising improvements in generating visually coherent and contextually rich movie scenes, paving the way for further exploration in cinematic video synthesis.

电影生成关键帧合成多角色交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。