用多智能体系统自动生成广告横幅,让设计更准更稳。
Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents
- 分层多模态智能体协作生成,支持多风格迭代优化。
- 仅需自然语言和logo图输入,自动修正多种设计错误。
- 适合广告设计自动化,尤其需要精准排版的品牌方。
近期生成模型如GPT-4o在高质量图像生成与文字渲染方面表现强劲。然而,商业广告横幅等设计任务不仅要求视觉保真,还需结构化布局、精准排版与品牌一致性。本文提出MIMO(Mirror In-the-Model)——一种面向自动广告横幅生成的智能体精炼框架。MIMO结合分层多模态代理系统(MIMO-Core)与协调循环(MIMO-Loop),探索多种风格方向并迭代提升设计质量。仅需自然语言提示与品牌标志图作为输入,MIMO即可自动检测并修正多种生成错误。实验表明,该方法在真实广告设计场景中显著优于现有扩散模型与LLM基线。
原文摘要 · Abstract (English)
Recent generative models such as GPT-4o have shown strong capabilities in producing high-quality images with accurate text rendering. However, commercial design tasks like advertising banners demand more than visual fidelity -- they require structured layouts, precise typography, consistent branding, and more. In this paper, we introduce MIMO (Mirror In-the-Model), an agentic refinement framework for automatic ad banner generation. MIMO combines a hierarchical multi-modal agent system (MIMO-Core) with a coordination loop (MIMO-Loop) that explores multiple stylistic directions and iteratively improves design quality. Requiring only a simple natural language based prompt and logo image as input, MIMO automatically detects and corrects multiple types of errors during generation. Experiments show that MIMO significantly outperforms existing diffusion and LLM-based baselines in real-world banner design scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。