无需训练,让生成故事中角色始终一致
Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

- 用初始描述信息补全后续提示中的角色细节
- 在多个基线模型上显著提升角色一致性
- 插件式设计,零成本适配主流扩散模型
视觉讲故事需要生成符合叙事且角色身份一致的图像。在自由形式的故事生成中,角色仅在首次出现时被完整描述,后续通过类型提及或代词指代。这种设定更贴近自然叙述,但后续提示可能遗漏关键身份语义,导致角色一致性难维持。我们提出 extbf{Sidecar},一个即插即用的语义增强模块,从初始描述中保留实体级信息,并将缺失语义注入后续提示嵌入。Sidecar 不需额外训练,也不改变基底扩散模型结构。在 FreeStoryBench 上的实验表明,Sidecar 在多个基于 SDXL 与 FLUX 的基线上持续提升提示-图像对齐与角色一致性,计算开销可忽略。
原文摘要 · Abstract (English)
Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introduced and is later referred to by a type-level mention or pronoun. Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain. We propose \textbf{Sidecar}, a plug-and-play semantic augmentation module that preserves entity-level information from the initial description and injects the missing semantics into later prompt embeddings. Sidecar requires no additional training and does not modify the architecture of the base diffusion model. Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。