arXiv:2412.07375cs.CV2024-12被引 12

用知识图谱增强角色一致性,让故事生成更精准

StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization

  • 构建角色知识图谱,统一建模人物与关系
  • 生成图像时角色身份一致,语义对齐提升13.44%
  • 适合需要多角色连贯视觉化的故事创作场景

故事可视化在人工智能领域受到越来越多关注。然而,现有方法在保持角色身份一致性和文本语义对齐之间仍难以平衡,主要由于缺乏对故事场景的精细语义建模。为此,我们提出一种新的知识图谱——角色图(Character Graph, CG),全面表征人物、人物属性及其相互关系等故事相关知识。在此基础上,我们设计了StoryWeaver图像生成模型,通过角色图(C-CG)实现角色定制化生成,能够生成语义丰富且角色一致的故事图像。为进一步提升多角色生成效果,我们在StoryWeaver中引入知识增强的空间引导(KE-SG),精准注入角色语义信息。通过新基准TBC-Bench上的大量实验验证,结果表明StoryWeaver不仅可生成生动的故事画面,还能在多种场景下准确传达角色身份,存储效率高,例如在DINO-I上平均提升+9.03%,在CLIP-T上提升+13.44%。消融实验进一步证明了各模块的有效性。代码与数据集已公开于https://github.com/Aria-Zhangjl/StoryWeaver。

原文摘要 · Abstract (English)

Story visualization has gained increasing attention in artificial intelligence. However, existing methods still struggle with maintaining a balance between character identity preservation and text-semantics alignment, largely due to a lack of detailed semantic modeling of the story scene. To tackle this challenge, we propose a novel knowledge graph, namely Character Graph (\textbf{CG}), which comprehensively represents various story-related knowledge, including the characters, the attributes related to characters, and the relationship between characters. We then introduce StoryWeaver, an image generator that achieve Customization via Character Graph (\textbf{C-CG}), capable of consistent story visualization with rich text semantics. To further improve the multi-character generation performance, we incorporate knowledge-enhanced spatial guidance (\textbf{KE-SG}) into StoryWeaver to precisely inject character semantics into generation. To validate the effectiveness of our proposed method, extensive experiments are conducted using a new benchmark called TBC-Bench. The experiments confirm that our StoryWeaver excels not only in creating vivid visual story plots but also in accurately conveying character identities across various scenarios with considerable storage efficiency, \emph{e.g.}, achieving an average increase of +9.03\% DINO-I and +13.44\% CLIP-T. Furthermore, ablation experiments are conducted to verify the superiority of the proposed module. Codes and datasets are released at https://github.com/Aria-Zhangjl/StoryWeaver.

故事生成知识图谱角色一致性图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。