让故事图像序列更合逻辑,避免动作脱节和叙事混乱。
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization

- 用多智能体系统显式建模角色、因果链和故事一致性。
- 在新基准LogicTale上显著提升叙事逻辑与视觉质量。
- 适合研究视觉叙事生成或需严谨逻辑的视频创作人群。
当前多模态系统在生成连贯且可传达的视觉序列(如图像序列、视频)方面仍面临重大挑战。尽管视觉质量与世界知识融合有所进步,现有模型仍难以维持逻辑连贯性,常导致动作断裂、叙事碎片化和主线模糊。我们将其归因于对视觉逻辑——即角色、动作与场景随时间的感知与因果一致性——的关注不足。为此,我们提出一个面向多图像故事可视化的逻辑感知框架LogiStory。该框架的核心创新在于显式建模视觉逻辑。通过设计多智能体系统,实现角色定位、因果链提取与故事级一致性验证,将叙事连贯性从图像生成的隐含结果转变为明确建模目标。该设计有效连接结构化故事规划与视觉生成,同时提升叙事清晰度与视觉质量。此外,我们构建了LogicTale基准,包含丰富标注的故事数据,强调因果推理与视觉逻辑可解释性,并建立自动与人工评估协议以衡量视觉逻辑与感知质量。实验表明,所提方法显著提升生成视觉故事的叙事逻辑。本工作为通用图像序列与视频生成任务中的视觉逻辑建模与强制奠定了基础。
原文摘要 · Abstract (English)
Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing models still struggle to maintain logical flow, often resulting in disjointed actions, fragmented narratives, and unclear storylines. We attribute these issues to the lack of attention to visual logic, a critical yet underexplored dimension of visual sequence generation that we define as the perceptual and causal coherence among characters, actions, and scenes over time. To bridge this gap, we propose a logic-aware multi-image story visualization framework, LogiStory. The framework is built around the central innovation of explicitly modeling visual logic in story visualization. To realize this idea, we design a multi-agent system that grounds roles, extracts causal chains, and verifies story-level consistency, transforming narrative coherence from an implicit byproduct of image generation into an explicit modeling objective. This design effectively bridges structured story planning with visual generation, enhancing both narrative clarity and visual quality in story visualization. Furthermore, to evaluate the generation capacity, we construct LogicTale, a benchmark comprising richly annotated stories, emphasizing causal reasoning, and visual logic interpretability. We establish comprehensive automatic and human evaluation protocols designed to measure both visual logic and perceptual quality. Experiments demonstrate that our approach significantly improves the narrative logic of generated visual stories. This work provides a foundational step towards modeling and enforcing visual logic in general image sequence and video generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。