arXiv:2606.23221cs.CVcs.AI2026-06被引 1

让AI图像生成学会主动提问和查资料,解决模糊指令难题

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

论文配图:RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
图 1 · 摘自论文原文
  • 设计闭环提问-求解机制,自动识别知识盲区并规划补救动作
  • 在两个基准测试中分别提升19.70和0.313分,达开源模型最优水平
  • 无需训练即可接入现有模型,适合需要精准理解指令的生成场景

近年来图像生成与编辑技术在遵循指令和视觉保真度方面取得显著进展。然而,在处理模糊意图、逻辑推理及分布外(OOD)知识时,现有模型因缺乏深度推理能力与实时外部信息支持,表现欠佳。尽管统一理解与生成模型试图弥合这一差距,仍受限于参数规模和静态知识。受智能体范式启发,我们提出RS-Gen:一个即插即用、免训练的多阶段图像生成智能体框架。该框架创新性引入“提问-求解”闭环机制,可准确识别逻辑问题与知识缺口,自主规划行动以填补信息不足并执行深层逻辑推理。大量实验表明,RS-Gen显著拓展了基础图像生成与编辑模型的能力边界。具体而言,在WISE Verified与RISEBench基准上,对Qwen-Image与Qwen-Image-Edit-2511分别带来0.313与19.70的绝对性能提升,成功将其推至开源模型中的最先进水平。

原文摘要 · Abstract (English)

Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD) knowledge, existing image models often yield sub-optimal results due to a lack of deep reasoning capabilities and real-time external information. Although emerging unified understanding-and-generation models attempt to bridge this gap, they remain constrained by their intrinsic parameter scales and static knowledge gaps. Inspired by agentic paradigms, we propose RS-Gen: a plug-and-play, training-free, multi-stage image agentic framework. RS-Gen innovatively introduces a "Questioning-and-Solving" closed-loop mechanism to accurately identify logical issues and knowledge gaps, autonomously planning actions to bridge information deficits and execute deep logical reasoning. Extensive experiments demonstrate that RS-Gen significantly expands the capability boundaries of foundational image generation and editing models. Specifically, on the WISE Verified and RISEBench benchmarks, RS-Gen yields substantial absolute performance gains of 0.313 for Qwen-Image and 19.70 for Qwen-Image-Edit-2511, respectively, successfully elevating both to the state-of-the-art (SOTA) level among open-source models.

图像生成智能体逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。