arXiv:2507.13285cs.CL2025-07

多智能体协作迭代生成高质量图文叙事,逻辑更连贯、布局更合理。

Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis

  • 用多智能体协同规划故事脉络,动态优化视觉布局。
  • 在多个数据集上超越基线模型,接近人工专家水平。
  • 适合需要自动创作专业级演示内容的团队使用。

自动生成高质量媒体演示面临挑战,需兼顾内容提取、叙事规划、视觉设计与整体质量优化。现有方法常出现逻辑不一致和布局不佳问题,难以达到专业标准。为此,我们提出RCPS(反思性连贯演示合成)框架,包含三大核心组件:(1) 深度结构化叙事规划;(2) 自适应布局生成;(3) 迭代优化循环。此外,我们提出PREVAL,一种基于偏好评估的框架,采用带理由增强的多维度模型,从内容、连贯性与设计三个维度评估演示质量。实验表明,RCPS在所有质量维度上均显著优于基线方法,生成结果接近人类专家水平。PREVAL与人工判断高度相关,验证其作为自动化评估工具的可靠性。

原文摘要 · Abstract (English)

Automated generation of high-quality media presentations is challenging, requiring robust content extraction, narrative planning, visual design, and overall quality optimization. Existing methods often produce presentations with logical inconsistencies and suboptimal layouts, thereby struggling to meet professional standards. To address these challenges, we introduce RCPS (Reflective Coherent Presentation Synthesis), a novel framework integrating three key components: (1) Deep Structured Narrative Planning; (2) Adaptive Layout Generation; (3) an Iterative Optimization Loop. Additionally, we propose PREVAL, a preference-based evaluation framework employing rationale-enhanced multi-dimensional models to assess presentation quality across Content, Coherence, and Design. Experimental results demonstrate that RCPS significantly outperforms baseline methods across all quality dimensions, producing presentations that closely approximate human expert standards. PREVAL shows strong correlation with human judgments, validating it as a reliable automated tool for assessing presentation quality.

多智能体叙事生成视觉设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。