arXiv:2603.17295cs.CVcs.AI2026-03

通过注意力机制与偏好优化,提升故事生成的连贯性与风格一致性。

Directing the Narrative: A Finetuning Method for Controlling Coherence and Style in Story Generation

  • 采用分组共享注意力机制,实现跨帧身份信息无损传递。
  • 在ViStoryBench上提升角色一致性和风格一致性,分别达+10.0和+18.7。
  • 无需外部编码器,适合需要长期视觉一致性的叙事生成任务。

故事可视化需生成与叙事演进语义一致且保持角色身份与视觉风格一致的序列图像。现有方法常面临主体不一致与身份漂移问题,尤其在复杂互动或长篇叙事中。为此,我们提出一种两阶段框架:首先引入分组共享注意力(GSA),在注意力层内实现跨样本无损信息流动,使模型在不依赖外部编码器的前提下结构化编码身份对应关系;其次采用直接偏好优化(DPO),基于整体偏好数据同时提升视觉保真度与身份保留能力,避免传统方法中冲突辅助损失的干扰。在ViStoryBench基准上的大量评估表明,该方法达到新SOTA,角色一致性(CIDS)提升+10.0,风格一致性(CSD)提升+18.7,同时保持高保真生成质量。

原文摘要 · Abstract (English)

Story visualization requires generating sequential imagery that aligns semantically with evolving narratives while maintaining rigorous consistency in character identity and visual style. However, existing methodologies often struggle with subject inconsistency and identity drift, particularly when depicting complex interactions or extended narrative arcs. To address these challenges, we propose a cohesive two-stage framework designed for robust and consistent story generation. First, we introduce Group-Shared Attention (GSA), a mechanism that fosters intrinsic consistency by enabling lossless cross-sample information flow within attention layers. This allows the model to structurally encode identity correspondence across frames without relying on external encoders. Second, we leverage Direct Preference Optimization (DPO) to align generated outputs with human aesthetic and narrative standards. Unlike conventional methods that rely on conflicting auxiliary losses, our approach simultaneously enhances visual fidelity and identity preservation by learning from holistic preference data. Extensive evaluations on the ViStoryBench benchmark demonstrate that our method establishes a new state-of-the-art, significantly outperforming strong baselines with gains of +10.0 in Character Identity (CIDS) and +18.7 in Style Consistency (CSD), all while preserving high-fidelity generation.

故事生成视觉一致性偏好优化注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。