用大模型生成能讲复杂故事的图像,让视觉更有逻辑层次。
Generating Storytelling Images with Rich Chains-of-Reasoning
- 先用大模型构思多层推理链条,再合成图像
- 新评估框架验证图像语义复杂度与图文一致率
- 小模型也能生成优质故事图像,适合资源有限场景
一张图像可通过逻辑关联的视觉线索讲述一个完整故事,形成多层次的推理链(Chains-of-Reasoning, CoRs)。我们定义这类语义丰富的图像为讲故事图像(Storytelling Images),其蕴含多层信息,激发主动解读,适用于插画、认知筛查等场景。然而,此类图像稀缺且难生成。为此,我们提出讲故事图像生成任务,并设计两阶段流程 StorytellingPainter:结合大语言模型(LLMs)的推理能力与文本到图像(T2I)生成技术。同时构建专用评估框架,衡量语义复杂度、多样性及图文对齐程度。针对故事生成的关键作用,我们引入轻量级 Mini-Storytellers,弥合小型模型与专有大模型间的性能差距。实验表明该方法可行有效。
原文摘要 · Abstract (English)
A single image can convey a compelling story through logically connected visual clues, forming Chains-of-Reasoning (CoRs). We define these semantically rich images as Storytelling Images. By conveying multi-layered information that inspires active interpretation, these images enable a wide range of applications, such as illustration and cognitive screening. Despite their potential, such images are scarce and complex to create. To address this, we introduce the Storytelling Image Generation task and propose StorytellingPainter, a two-stage pipeline combining the reasoning of Large Language Models (LLMs) with Text-to-Image (T2I) synthesis. We also develop a dedicated evaluation framework assessing semantic complexity, diversity, and text-image alignment. Furthermore, given the critical role of story generation in the task, we introduce lightweight Mini-Storytellers to bridge the performance gap between small-scale and proprietary LLMs. Experimental results demonstrate the feasibility of our approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。