arXiv:2509.22940cs.CLcs.CV2025-09EMNLP被引 1

用大模型生成故事场景图,构建了首个跨模态叙事数据集。

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

  • 用大模型解析故事文本,生成图像提示词驱动文生图
  • 构建包含3000+场景的SceneIllustrations数据集,支持质量评估
  • 验证大模型能有效提取隐含场景信息,提升图文生成效果

生成式AI使得内容跨媒介转换变得便捷,尤其在叙事场景中,视觉插图可生动呈现文字故事。本文聚焦叙事场景插图任务,提出一种基于大模型(LLM)的文生图提示生成流程:将原始故事文本输入大模型,由其生成适配文生图模型的描述性提示,进而生成对应场景图像。该方法应用于一个主流故事语料库,合成数千个故事场景的插图。通过人工标注进行成对质量评估,构建出名为SceneIllustrations的新数据集,以供未来跨模态叙事研究使用。分析表明,大模型能有效提取故事文本中隐含的场景知识;该能力显著提升了插图生成与评价的质量。

原文摘要 · Abstract (English)

Generative AI has established the opportunity to readily transform content from one medium to another. This capability is especially powerful for storytelling, where visual illustrations can illuminate a story originally expressed in text. In this paper, we focus on the task of narrative scene illustration, which involves automatically generating an image depicting a scene in a story. Motivated by recent progress on text-to-image models, we consider a pipeline that uses LLMs as an interface for prompting text-to-image models to generate scene illustrations given raw story text. We apply variations of this pipeline to a prominent story corpus in order to synthesize illustrations for scenes in these stories. We conduct a human annotation task to obtain pairwise quality judgments for these illustrations. The outcome of this process is the SceneIllustrations dataset, which we release as a new resource for future work on cross-modal narrative transformation. Through our analysis of this dataset and experiments modeling illustration quality, we demonstrate that LLMs can effectively verbalize scene knowledge implicitly evoked by story text. Moreover, this capability is impactful for generating and evaluating illustrations.

叙事生成文生图大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。