让3D分镜自动保持角色一致且可精确编辑,解决传统方法难用、生成效果不稳的痛点。
StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
- 用连续记忆图分离全局资产与镜头变量,确保长序列一致性
- 在统一坐标系中实例化角色,视觉身份稳定无漂移
- 支持相机与道具直接编辑,适合影视动画快速原型设计
分镜是电影、动画和游戏视觉叙事的核心技能。但自动化分镜需同时满足跨镜头一致性和显式可编辑性,当前方法极少能兼顾。2D扩散模型虽生成画面生动,却常出现角色身份漂移且几何控制有限;传统3D动画流程虽一致且可编辑,但依赖专家、耗时费力。我们提出StoryBlender,一个基于故事中心反思机制的3D分镜生成框架。其三阶段流程包括:(1) 语义-空间锚定,构建连续性记忆图,解耦全局资产与镜头特异性变量以实现长程一致性;(2) 标准化资产生成,在统一坐标空间实例化实体以维持视觉身份;(3) 空间-时间动力学,通过视觉度量实现布局设计与镜头演进。通过层级代理在验证循环中协作,系统利用引擎反馈迭代修正空间幻觉。最终生成的原生3D场景支持对相机与视觉资产的直接精准编辑,同时保持多镜头间无懈可击的一致性。实验表明,StoryBlender在一致性和可编辑性上显著优于扩散基与3D基基线。代码、数据与演示视频将发布于 https://engineeringai-lab.github.io/StoryBlender/
原文摘要 · Abstract (English)
Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter-shot consistency and explicit editability. While 2D diffusion-based generators produce vivid imagery, they often suffer from identity drift along with limited geometric control; conversely, traditional 3D animation workflows are consistent and editable but require expert-heavy, labor-intensive authoring. We present StoryBlender, a grounded 3D storyboard generation framework governed by a Story-centric Reflection Scheme. At its core, we propose the StoryBlender system, which is built on a three-stage pipeline: (1) Semantic-Spatial Grounding, to construct a continuity memory graph to decouple global assets from shot-specific variables for long-horizon consistency; (2) Canonical Asset Materialization, to instantiate entities in a unified coordinate space to maintain visual identity; and (3) Spatial-Temporal Dynamics, to achieve layout design and cinematic evolution through visual metrics. By orchestrating multiple agents in a hierarchical manner within a verification loop, StoryBlender iteratively self-corrects spatial hallucinations via engine-verified feedback. The resulting native 3D scenes support direct, precise editing of cameras and visual assets while preserving unwavering multi-shot continuity. Experiments demonstrate that StoryBlender significantly improves consistency and editability over both diffusion-based and 3D-grounded baselines. Code, data, and demonstration video will be available on https://engineeringai-lab.github.io/StoryBlender/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。