arXiv:2509.04123cs.CV2025-09被引 3

让多角色故事生成更连贯,对话准确匹配角色

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

  • 用提示学习生成每帧描述与对话,结合注意力控制角色互动
  • 通过身份一致自注意力提升跨帧角色一致性,减少画面错乱
  • 对话气泡精准定位并分配给对应角色,适合剧情可视化应用

文本到故事的视觉化因需保持多角色跨帧一致互动而极具挑战。现有方法在角色一致性上表现不佳,导致生成伪影和对话错位,影响叙事连贯性。为此,我们提出TaleDiffusion框架,采用迭代生成流程,通过预训练大模型进行上下文学习,生成每帧的描述、角色细节和对话;利用基于边界注意力的逐框掩码技术控制角色交互,降低伪影;引入身份一致自注意力机制确保跨帧角色一致性,结合区域感知交叉注意力实现物体精确定位;对话则通过CLIPSeg生成气泡并准确分配给对应角色。实验表明,TaleDiffusion在一致性、噪声抑制和对话渲染方面均优于现有方法。

原文摘要 · Abstract (English)

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inaccurate dialogue rendering, which results in disjointed storytelling. In response, we introduce TaleDiffusion, a novel framework for generating multi-character stories with an iterative process, maintaining character consistency, and accurate dialogue assignment via postprocessing. Given a story, we use a pre-trained LLM to generate per-frame descriptions, character details, and dialogues via in-context learning, followed by a bounded attention-based per-box mask technique to control character interactions and minimize artifacts. We then apply an identity-consistent self-attention mechanism to ensure character consistency across frames and region-aware cross-attention for precise object placement. Dialogues are also rendered as bubbles and assigned to characters via CLIPSeg. Experimental results demonstrate that TaleDiffusion outperforms existing methods in consistency, noise reduction, and dialogue rendering.

故事生成多角色对话渲染图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。