arXiv:2606.25079cs.CV2026-06被引 1

无需训练即可实现自由叙事下的角色一致性生成

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

论文配图:FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling
图 1 · 摘自论文原文
  • 通过实体锚定特征复用,解决自由形式提示中的角色一致问题
  • 在自由叙事下角色一致性指标提升18.7%,优于现有方法
  • 适合需要自然对话式描述的视频/图像故事生成场景

视觉讲故事旨在生成与叙事提示一致且角色外观保持一致的图像序列。现有免训练方法通过复用注意力特征提升角色一致性,但依赖结构化提示——即每条提示中重复完整角色描述。这一假设虽简化任务,却偏离自然叙事:角色通常仅首次出现时被完整描述,后续使用代词或类型表达。本文提出【FreeStory】,一种免训练框架,将自由形式提示下的角色一致性重构为实体锚定特征复用。该方法将引用提及与对应角色描述关联,并结合动态角色掩码、对应感知特征匹配、键值注入和查询混合,在保留生成多样性的同时维持角色身份。我们还构建了【FreeStoryBench】基准,涵盖单角色与多角色故事。实验表明,FreeStory在结构化基准上达到免训练方法最优性能,并在自由形式提示下显著优于基线,整体一致性更强。

原文摘要 · Abstract (English)

Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images. Recent training-free methods improve character consistency by reusing attention features, but rely on structured prompts where full character descriptions are repeated in every prompt. This assumption simplifies the task but deviates from natural storytelling, where characters are typically introduced once and later referred to using pronouns or type-based expressions. We propose \textbf{FreeStory}, a training-free framework that reformulates character consistency under free-form prompts as entity-grounded feature reuse. Our method associates reference mentions with their corresponding character descriptions and combines dynamic character masks, correspondence-aware feature matching, key-value injection, and query blending to preserve identity while retaining generation diversity. We also introduce \textbf{FreeStoryBench}, a benchmark for this setting that includes both single- and multi-character stories. Experiments show that FreeStory achieves state-of-the-art performance among training-free methods on structured benchmarks and stronger overall consistency over baselines under free-form prompts.

视觉讲故事角色一致性免训练自由叙事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。