arXiv:2508.09983cs.CVcs.GR2025-08被引 15

无需训练,让扩散模型生成连贯有叙事张力的分镜图

Story2Board: A Training-Free Approach for Expressive Storyboard Generation

  • 用轻量级一致性机制保持角色跨画面一致
  • 在不修改架构下提升分镜布局多样性与背景连贯性
  • 适合想快速生成叙事性强分镜的创作者

我们提出Story2Board,一种从自然语言生成富有表现力分镜的无训练框架。现有方法仅关注主体身份,忽视空间构图、背景演变和叙事节奏等关键视觉叙事要素。为此,我们设计了一个轻量级一致性框架,包含两个组件:潜空间面板锚定(Latent Panel Anchoring),用于跨画面保持角色参考一致性;互注意力值混合(Reciprocal Attention Value Mixing),通过强互注意力的令牌对软融合视觉特征。二者协同增强画面连贯性,无需模型结构调整或微调,即可使当前最先进的扩散模型生成视觉多样且一致的分镜。为结构化生成过程,我们采用现成语言模型将自由文本故事转化为带场景提示的分镜级指令。为评估,我们构建了丰富的分镜基准(Rich Storyboard Benchmark),涵盖开放域叙事,用于评估布局多样性、背景相关叙事与一致性;并引入新指标“场景多样性”以量化分镜间空间与姿态变化。定性与定量结果及用户研究均表明,Story2Board生成的分镜更具动态性、连贯性和叙事吸引力,优于现有基线。

原文摘要 · Abstract (English)

We present Story2Board, a training-free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing. To address this, we introduce a lightweight consistency framework composed of two components: Latent Panel Anchoring, which preserves a shared character reference across panels, and Reciprocal Attention Value Mixing, which softly blends visual features between token pairs with strong reciprocal attention. Together, these mechanisms enhance coherence without architectural changes or fine-tuning, enabling state-of-the-art diffusion models to generate visually diverse yet consistent storyboards. To structure generation, we use an off-the-shelf language model to convert free-form stories into grounded panel-level prompts. To evaluate, we propose the Rich Storyboard Benchmark, a suite of open-domain narratives designed to assess layout diversity and background-grounded storytelling, in addition to consistency. We also introduce a new Scene Diversity metric that quantifies spatial and pose variation across storyboards. Our qualitative and quantitative results, as well as a user study, show that Story2Board produces more dynamic, coherent, and narratively engaging storyboards than existing baselines.

分镜生成扩散模型叙事理解无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。