用物理世界结构化叙事取代像素采样,实现视频生成的精准控制。
World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

- 将视频生成拆解为世界构建与像素渲染两阶段,用4D物理世界表征替代黑箱采样
- 支持文字、视频、草图等多模态输入,生成可编辑的几何、运动、镜头和光照结构
- 显著减少随机生成次数,让专业创作者能精确控制画面布局与影视效果
现有视频生成模型将视频视为像素分布采样问题,忽视了实例级4D(3D+时间)物理世界的显式建模,导致内容创作者无法以确定性方式指定几何、运动、相机参数或光照,造成低效且昂贵的“抽卡”循环。为此,我们提出世界叙事模型(WNM),将要渲染的内容(结构化物理叙事)与渲染方式(像素生成过程)解耦。WNM以协同智能体将文本、参考视频和草图等稀疏多模态输入转化为可编辑的4D世界表示,包含场景几何、物体布局、角色/动物骨骼运动、轨迹、相机运动和光照,具备定量且物理意义明确的粒度。该表示作为确定性结构蓝图,驱动现有视频基础模型(冻结或轻量微调)生成最终画面,使基础模型成为忠实的神经着色器。基于此引擎,我们的平台支持与专业电影制作流程对齐的自动世界生成与预可视化,导演控制台实现无缝人工修正。实验表明,WNM大幅减少概率性‘gacha’调用,生成视频的布局、运动与摄影风格高度契合创作意图。框架开放模块化,允许世界表征、控制代理和适配器等组件独立优化。项目网站:https://glassroom.sjtu.edu.cn/WNM/
原文摘要 · Abstract (English)
The fundamental obstacle to industrial grade video generation is the lack of controllability: existing models treat video as a pixel distribution sampling problem, bypassing the explicit, instance level $4D$ $(3D + T)$ physical world. Consequently, content creators cannot specify geometry, motion, camera parameters, or lighting in a deterministic, quantitative way, leading to the infamous ''gacha'' loop that makes professional content creation prohibitively inefficient and expensive. To address this, we introduce the World Narrative Model (WNM), a paradigm that decouples what to render -- the structured physical narrative -- from how to render -- the pixel generation process. WNM replaces end-to-end black-box sampling with orchestrated $4D$ pre-visualization for media generation. Collaborative agents translate sparse multimodal inputs, including text, reference videos, and sketches, into a fully editable world representation with scene geometry, object layouts, character/animal skeleton motion, trajectories, camera motion, and lighting at quantitative, physically meaningful granularity. This representation acts as a deterministic structural blueprint that drives existing video foundation models, either frozen or lightly adapted, to render final footage, turning the base model into a faithful neural shader. Built on this engine, our human-AI platform supports automatic world generation and pre-visualization aligned with professional filmmaking pipelines, while director consoles enable seamless human refinement. Experiments show that WNM greatly reduces probabilistic ``gacha'' calls and produces videos whose layout, motion, and cinematography closely follow creator intent. The framework is open and modular, allowing each component, such as world representation, control agents, and adapters, to be independently improved. Project website: https://glassroom.sjtu.edu.cn/WNM/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。