用分层智能体协调图文生成,让网页设计更统一美观。
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

- 分层规划+自我反思,统筹全局布局与局部元素生成
- 在多模态网页生成任务中超越现有基线模型表现
- 适合需要高质量、一致视觉风格的网页自动化设计人群
人工智能生成内容(AIGC)工具使图像、视频和可视化内容可按需生成,为现代UI/UX设计提供了灵活且日益普及的新范式。然而,直接将这些工具集成到自动化网页生成中常导致风格不一致和整体连贯性差,因元素独立生成所致。我们提出MM-WebAgent,一种用于多模态网页生成的分层智能体框架,通过分层规划与迭代自我反思来协调基于AIGC的元素生成。该框架联合优化全局布局、局部多模态内容及其整合,生成视觉一致且结构连贯的网页。我们进一步构建了一个多模态网页生成基准和多层级评估协议,实现系统性评估。实验表明,MM-WebAgent在多模态元素生成与整合方面显著优于代码生成与基于智能体的基线模型。代码与数据:https://aka.ms/mm-webagent。
原文摘要 · Abstract (English)
The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for webpage design, offering a flexible and increasingly adopted paradigm for modern UI/UX. However, directly integrating such tools into automated webpage generation often leads to style inconsistency and poor global coherence, as elements are generated in isolation. We propose MM-WebAgent, a hierarchical agentic framework for multimodal webpage generation that coordinates AIGC-based element generation through hierarchical planning and iterative self-reflection. MM-WebAgent jointly optimizes global layout, local multimodal content, and their integration, producing coherent and visually consistent webpages. We further introduce a benchmark for multimodal webpage generation and a multi-level evaluation protocol for systematic assessment. Experiments demonstrate that MM-WebAgent outperforms code-generation and agent-based baselines, especially on multimodal element generation and integration. Code & Data: https://aka.ms/mm-webagent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。