arXiv:2609.05171cs.CV2026-09

让图像生成更懂常识,靠多模态智能体自动查证与整合信息。

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

论文配图:WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing
图 1 · 摘自论文原文
  • 用多模态运行时管理证据,实时整合图文信息。
  • 构建23000条带验证的生成轨迹,提升复杂任务可靠性。
  • 适合需要跨模态推理的图像生成与编辑研究者使用。

图像生成与编辑模型虽快速进步,但在涉及外部世界知识的提示下仍不可靠。有限的参数化知识难以支持直接或先推理后生成的方法获取必要事实与视觉特征。现有智能体方法虽借助检索工具缓解此问题,但仍受限于视觉验证不足、策略模型过载及检索到的文本与视觉证据融合薄弱。为此,我们提出WeAgent-MMGenEdit,一套包含多模态运行时、可扩展数据构建流水线、全面基准测试和后训练方法的完整解决方案。首先引入WeAgent-Harness,一个具备持续证据管理与专用验证/整合工具的多模态运行时,将检索到的多模态证据组织成密集载体。在此基础上,开发了用于提示合成与智能体轨迹收集的可扩展流水线,生成23,000条监督轨迹与14,700个强化学习任务,并配备三层可验证检查清单。进一步提出WeBench-MMGenEdit,一个涵盖知识密集型图像生成与多图编辑的双语基准。最后,基于SFT与强化学习的双向后训练方案优化智能体策略与图像后端。整体系统使300亿总参数/30亿活跃参数的策略模型优于同规模模型,逼近1万亿参数智能体表现。

原文摘要 · Abstract (English)

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approaches from recovering the required facts and visual appearances. Existing agentic generation and editing methods mitigate this limitation with retrieval tools, yet remain constrained by insufficient visual verification, overloaded policy models, and weak integration of retrieved textual and visual evidence. To address these limitations, we present WeAgent-MMGenEdit, a full-stack recipe including a multimodal harness, a scalable data construction pipeline, a comprehensive benchmark, and post-training methods for the agent policy and image backend. We first introduce WeAgent-Harness, a multimodal runtime with persistent evidence management and dedicated verification and integration tools that organize retrieved multimodal evidence into a dense carrier. Upon this, we develop a scalable pipeline for prompt synthesis and agentic trajectory collection, yielding 23K supervised trajectories and 14.7K RL tasks with three-layer verifiable checklists. We further introduce WeBench-MMGenEdit, a bilingual benchmark covering both knowledge-intensive image generation and multi-image editing. Finally, a two-sided post-training recipe based on SFT and RL improves the agent policy and image backend. Together, WeAgent-MMGenEdit enables a 30B-total/3B-active policy to outperform similarly sized policy models and approach the performance of a 1T-parameter agent.

图像生成智能体多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。