arXiv:2602.01756cs.CV2026-02被引 15

让图像生成像人一样思考:先推理再创作

Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation

  • 引入智能体式思维流程,边搜索边推理生成图像
  • 在新概念和复杂推理任务上,基线模型能力实现零到一突破
  • 适合需要理解隐含意图与动态知识的生成场景

尽管文本到图像生成已达到前所未有的保真度,但现有模型大多仅作为静态的文本到像素解码器,难以理解用户隐含意图。虽有统一理解和生成模型提升了意图识别能力,但在复杂知识推理任务中仍受限于单一模型架构。此外,受制于静态内部先验,这些模型无法适应现实世界的动态变化。为此,我们提出Mind-Brush,一个统一的智能体框架,将生成过程转变为动态、知识驱动的工作流。模拟人类‘思考-调研-创作’模式,Mind-Brush主动检索多模态证据以支撑分布外概念,并使用推理工具解决隐含视觉约束。为严格评估该能力,我们构建了涵盖500个样本的Mind-Bench基准,覆盖实时新闻、新兴概念及数学与地理推理等域。大量实验表明,Mind-Brush显著提升统一模型能力,在Mind-Bench上使Qwen-Image基线实现零到一的能力跃升,同时在WISE和RISE等已有基准上表现更优。

原文摘要 · Abstract (English)

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user intentions. Although emerging unified understanding-generation models have improved intent comprehension, they still struggle to accomplish tasks involving complex knowledge reasoning within a single model. Moreover, constrained by static internal priors, these models remain unable to adapt to the evolving dynamics of the real world. To bridge these gaps, we introduce Mind-Brush, a unified agentic framework that transforms generation into a dynamic, knowledge-driven workflow. Simulating a human-like 'think-research-create' paradigm, Mind-Brush actively retrieves multimodal evidence to ground out-of-distribution concepts and employs reasoning tools to resolve implicit visual constraints. To rigorously evaluate these capabilities, we propose Mind-Bench, a comprehensive benchmark comprising 500 distinct samples spanning real-time news, emerging concepts, and domains such as mathematical and Geo-Reasoning. Extensive experiments demonstrate that Mind-Brush significantly enhances the capabilities of unified models, realizing a zero-to-one capability leap for the Qwen-Image baseline on Mind-Bench, while achieving superior results on established benchmarks like WISE and RISE.

图像生成智能体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。