让故事书生成更安全可靠,多智能体协同规划与校验
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

- 用多个智能体协作完成故事创作、脚本、绘图与全局修复
- 在角色身份和叙事逻辑上实现端到端一致性验证与修正
- 专为儿童安全设计,支持故事与图像的全程合规性检查
大型生成模型的进展推动了多模态生成的发展,但图文故事书生成仍面临挑战:现有方法多采用分阶段处理,难以实现整体多模态对齐。此外,尽管文本或图像生成中的安全对齐已有研究,但针对儿童安全约束的叙事规划与序列级多模态验证仍鲜有探索。为此,我们提出BookAgent,一个面向高质量、安全感知的视觉叙事生成的多智能体协同框架。不同于固定故事序列的生成模型,BookAgent从用户草稿出发,端到端完成故事书合成,联合规划、撰写脚本、生成图像并全局修复不一致问题。通过动态校准每页文字与视觉布局的对齐,以及在时间维度上验证并修正角色身份与叙事逻辑的全局不一致,实现精准多模态对齐。大量实验表明,BookAgent在叙事连贯性、视觉一致性与安全合规性方面显著优于现有方法,为复杂多模态创作提供可靠范式。代码将公开发布于 https://github.com/bogao-code/BookAgent/tree/main。
原文摘要 · Abstract (English)
Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an open challenge, where prior works mainly decompose this task into separate stages, and thus, holistic multi-modal grounding remains limited. Besides, while safety alignment is studied for text- or image-only generation, existing works rarely integrate child-specific safety constraints into narrative planning and sequence-level multi-modal verification. To address these limitations, we propose BookAgent, a safety-aware multi-agent collaboration framework designed for high-quality, safety-aware visual narratives. Different from prior story visualization models that assume a fixed storyline sequence, BookAgent targets end-to-end storybook synthesis from a user draft by jointly planning, scripting, illustrating, and globally repairing inconsistencies. To ensure precise multi-modal grounding, BookAgent dynamically calibrates page-level alignment between textual scripts and visual layouts. Furthermore, BookAgent calibrates holistic consistency from the temporal dimension, by verifying-then-rectifying global inconsistencies in character identity and storytelling logic. Extensive experiments demonstrate that BookAgent significantly outperforms current methods in narrative coherence, visual consistency, and safety compliance, offering a robust paradigm for reliable agents in complex multi-modal creation. The implementation will be publicly released at https://github.com/bogao-code/BookAgent/tree/main.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。