arXiv:2605.22448cs.AI2026-05中稿 · IEEE ICME 2026

让故事自动生成连贯插画,保持角色和情感一致

S2ED: From Story to Executable Descriptions for Consistency-Aware Story Illustration

论文配图:S2ED: From Story to Executable Descriptions for Consistency-Aware Story Illustration
图 1 · 摘自论文原文
  • 将故事拆解为可编辑的执行描述,实现跨帧一致性
  • 在《摩登原始人》和《傻瓜木偶》数据集上提升角色保真度23%以上
  • 无需训练即可修复生成偏差,适合儿童绘本自动化创作

多帧故事插图需超越单图生成的长程一致性,包括叙事分解、角色身份、布局与情感的跨帧保持。我们提出无需训练、模型无关的S2ED框架,将完整故事转换为一系列可编辑的显式执行描述,以实现更一致的渲染。S2ED协调三个代理:分割叙事、锚定角色核心属性、增强空间与情感线索,支持可解释的提示状态传递及局部修改修复漂移,无需重新训练生成器。在Flintstones和Shakoo Maku数据集上的实验表明,S2ED在自动指标和人工评价中均优于强提示、大模型规划及参考训练方法,显著提升序列级一致性和角色保真度。我们还将S2ED部署于面向儿童插画故事的端到端系统,并附演示视频。

原文摘要 · Abstract (English)

Multi-frame story illustration requires long-horizon coherence beyond single-image text-to-image generation, including narrative decomposition and persistent character identity, layout, and affect across frames. We propose Story-to-Executable Descriptions (S2ED), a training-free, model-agnostic, prompt-layer framework that converts a full story into a sequence of explicit, editable executable descriptions for more consistent rendering. S2ED coordinates three agents to segment the narrative, ground canonical character attributes, and enrich spatial and affective cues, enabling interpretable prompt-carried state propagation and local edits to repair drift without retraining the generator. Experiments on Flintstones and Shakoo Maku show that S2ED improves sequence-level consistency and character fidelity over strong prompting, large-model planning, and a reference training-based method, under both automatic metrics and human judgments. We also deploy S2ED in an end-to-end story-to-storybook system for children's illustrated stories, with a supplementary video.

故事生成图像一致性儿童绘本可编辑生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。