arXiv:2607.19038cs.CVcs.AI2026-07

将小说转为连贯电影,通过动态世界建模实现跨场景因果一致。

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

论文配图:FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
图 1 · 摘自论文原文
  • 分两阶段:先构建具状态的电影世界,再动态演化实体以保持剧情一致。
  • 在15部小说上测试,生成视频在叙事忠实度和跨镜头一致性上显著领先。
  • 适合影视生成、创意智能、多智能体系统研究者关注。

将小说转化为电影对生成式AI提出重大挑战,需将抽象文学叙述转化为长时序、多场景的视觉叙事。现有视频生成模型擅长短时序、单一场景内容,而小说到电影生成需在多样化场景中生成长时序内容,并支持动态实体状态演变。为此,我们将该任务形式化为动态电影世界建模,分为两个阶段:构建阶段将模糊的文学叙述具体化为具状态、可持久的世界实体;演化阶段则管理实体在剧情推进下的动态更新,确保跨场景因果一致性。我们提出FilmWorld,一个端到端的代理系统,由两组专用代理协作完成:构建侧代理执行结构化翻译、带视觉锚定的状态建模与状态驱动的镜头规划,逐步将文本投影为电影蓝图;演化侧代理执行状态锚定的视觉生成、跨镜头状态传播与闭环状态验证,维持因果一致与视觉连贯性。为弥补长序列生成评估空白,我们引入FilmEval,一个包含15个代表性小说的分级难度基准与九项自动化指标组成的评估框架,涵盖电影表现、影片一致性与小说忠实度三个维度。实验表明,FilmWorld持续优于当前最先进视频生成代理系统,尤其在叙事忠实度与跨场景一致性上提升明显。

原文摘要 · Abstract (English)

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entity states. To address this, we formalize novel-to-film generation as dynamic cinematic world modeling, decomposed into two phases: construction, which grounds abstract, underspecified literary narratives into concrete, stateful, and persistent world entities; and evolution, which governs how these entities dynamically update under plot progression to maintain causal consistency across scenes. We propose FilmWorld, an end-to-end agentic system where two groups of specialized agents collaborate to instantiate these phases. Construction-side agents perform narrative structured translation, world entity state modeling with visual anchoring, and state-driven shot planning, progressively projecting literary language into a cinematic blueprint. Evolution-side agents perform state-anchored visual generation, cross-shot dynamic state propagation, and closed-loop state verification to maintain causal consistency and visual coherence. To address the evaluation gap in long-form generation, we introduce FilmEval, a systematic evaluation framework that couples a difficulty-graded benchmark of 15 representative novels with an automated protocol of nine objective metrics spanning three dimensions: cinematic presentation, film consistency, and novel fidelity. Experiments demonstrate that FilmWorld consistently outperforms state-of-the-art video generation agent systems, with particularly pronounced improvements in narrative fidelity and cross-scene consistency.

视频生成多智能体小说转电影世界建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。