arXiv:2503.02399cs.CVcs.AI2025-03中稿 · ICASSP 2025被引 5

不依赖训练的多智能体框架,让故事图像更忠实于原意

VisAgent: Narrative-Preserving Story Visualization Framework

  • 用多个专业智能体协作解析故事结构并优化生成提示
  • 生成图像能更好保留叙事核心,避免关键信息丢失
  • 适合需要精准还原故事场景的创作与教育应用

故事可视化是将叙事元素转化为图像序列的过程。现有研究主要关注视觉上下文一致性,却常忽视故事的深层叙事本质,导致生成图像难以完整传达原有意图与细节。为此,我们提出VisAgent——一种无需训练的多智能体框架,用于理解并可视化给定故事中的关键场景。该框架通过故事提炼、语义一致性和上下文连贯性,采用智能体工作流:多个专用智能体协同完成(i)根据叙事结构优化分层提示,以及(ii)将优化后的提示、场景元素和主体位置等生成内容无缝融合至最终图像。实证验证表明,该框架在实际故事可视化应用中具有显著有效性。

原文摘要 · Abstract (English)

Story visualization is the transformation of narrative elements into image sequences. While existing research has primarily focused on visual contextual coherence, the deeper narrative essence of stories often remains overlooked. This limitation hinders the practical application of these approaches, as generated images frequently fail to capture the intended meaning and nuances of the narrative fully. To address these challenges, we propose VisAgent, a training-free multi-agent framework designed to comprehend and visualize pivotal scenes within a given story. By considering story distillation, semantic consistency, and contextual coherence, VisAgent employs an agentic workflow. In this workflow, multiple specialized agents collaborate to: (i) refine layered prompts based on the narrative structure and (ii) seamlessly integrate \gt{generated} elements, including refined prompts, scene elements, and subject placement, into the final image. The empirically validated effectiveness confirms the framework's suitability for practical story visualization applications.

故事生成多智能体图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。