让AI自动规划产品故事,生成连贯美观的多格海报
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

- 用自反思框架构建产品叙事结构,统一视觉风格
- 生成结果在叙事丰富度和视觉一致性上显著优于基线
- 适合电商、广告等需要高质量产品视觉呈现的场景
以叙事驱动的产品摄影已成为现代营销的主流范式,连贯的视觉叙事有助于传达产品价值并激发消费者情感共鸣。然而,现有图像生成方法缺乏结构化叙事规划与跨面板协调能力,常导致叙事薄弱和视觉不一致。实践中,叙事性产品摄影多以多网格拼贴形式呈现,多个视角或场景共同传递产品故事。为确保网格间视觉一致性与整体构图美感,我们以单幅统一图像生成拼贴,而非独立合成各面板。提出一种自反思智能体框架,给定产品主图及名称后,首先构建显式包含产品身份、使用情境与环境的叙事框架,并转化为共享视觉风格的互补网格。通过约束感知提示输入生成模型,联合合成拼贴图像。输出在内容合理性与摄影质量上进行评估,设有明确通过/修正判断机制。若评估失败,系统进行错误归因并实施针对性优化,实现通过迭代自我反思的持续改进。实验表明,该框架在美学质量、叙事丰富度与视觉连贯性方面均显著优于直接提示基线。
原文摘要 · Abstract (English)
Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engagement with consumers. However, existing image generation methods do not support structured narrative planning or cross-panel coordination, often resulting in weak storytelling and visual incoherence. In practice, narrative product photography is commonly presented as multi-grid collages, where multiple views or scenes jointly communicate a product narrative. To ensure visual consistency across grids and aesthetic harmony of the overall composition, we generate the collage as a single unified image rather than composing independently synthesized panels. We propose a self-reasoning agentic framework for narrative product grid collage generation. Given a product packshot and its name, the system first constructs a Product Narrative Framework that explicitly represents the product's identity, usage context, and situational environment, and translates it into complementary grids governed by a shared visual style. Constraint-aware prompts are then compiled and fed to a generation model that synthesizes the collage jointly. The generated output is evaluated on both content validity and photography quality, with explicit gates determining whether to proceed or refine. When evaluation fails, the system performs failure attribution and applies targeted refinement, enabling progressive improvement through iterative self-reflection. Experiments demonstrate that our framework consistently improves aesthetic quality, narrative richness, and visual coherence, compared to direct prompting baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。