用智能体框架提升图像生成对真实世界知识的依赖能力。
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
- 将图像生成拆解为理解、搜证、重描述、合成四步智能体流程。
- 在143K条轨迹上训练,实现在长尾概念上的生成性能显著提升。
- 适合需要外部知识推理的开放世界图像生成任务使用。
统一多模态模型虽能理解复杂现实知识并生成高质量图像,但仍依赖冻结的参数化知识,难以处理长尾和知识密集型概念。受智能体在真实任务中成功启发,我们提出 Unify-Agent——一种面向世界接地图像生成的统一多模态智能体。该模型将图像生成重构为包含提示理解、多模态证据搜索、接地重描述与最终合成的智能体流水线。为此,我们构建了定制化的多模态数据管道,标注了143,000条高质量智能体轨迹,以实现对完整生成过程的有效监督。我们还引入FactIP基准,覆盖12类具有文化意义和长尾特征的事实概念,明确要求外部知识接地。大量实验表明,Unify-Agent在多个基准和真实场景生成任务中显著优于其基础模型,接近最强闭源模型的世界知识能力。作为对基于智能体建模在世界接地图像生成中的早期探索,本工作凸显了推理、搜索与生成紧密耦合对于可靠开放世界智能体图像生成的价值。
原文摘要 · Abstract (English)
Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen parametric knowledge, which makes them struggle with real-world image generation involving long-tail and knowledge-intensive concepts. Inspired by the broad success of agents on real-world tasks, we explore agentic modeling to address this limitation. Specifically, we present Unify-Agent, a unified multimodal agent for world-grounded image synthesis, which reframes image generation as an agentic pipeline consisting of prompt understanding, multimodal evidence searching, grounded recaptioning, and final synthesis. To train our model, we construct a tailored multimodal data pipeline and curate 143K high-quality agent trajectories for world-grounded image synthesis, enabling effective supervision over the full agentic generation process. We further introduce FactIP, a benchmark covering 12 categories of culturally significant and long-tail factual concepts that explicitly requires external knowledge grounding. Extensive experiments show that our proposed Unify-Agent substantially improves over its base unified model across diverse benchmarks and real world generation tasks, while approaching the world knowledge capabilities of the strongest closed-source models. As an early exploration of agent-based modeling for world-grounded image synthesis, our work highlights the value of tightly coupling reasoning, searching, and generation for reliable open-world agentic image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。