让AI理解模糊图像需求,自动补全上下文生成精准图像。
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

- 用智能体框架逐步补全用户指令中的缺失信息
- 在多个评测中表现优于现有方法,达当前最优水平
- 适合需要理解复杂或不完整指令的图像生成场景
尽管文本到图像(T2I)模型取得显著进展,但在处理现实世界中常被省略、隐含或依赖最新知识的请求时仍存在困难。我们将其归因于‘上下文差距’:用户意图与T2I模型所需充分生成上下文之间的不匹配。为此,提出Qwen-Image-Agent,一种统一的智能体框架,以上下文为中心整合规划、推理、搜索、记忆和反馈机制。该框架将用户输入视为部分上下文,通过上下文感知规划与上下文锚定逐步构建完整生成上下文。其中,上下文感知规划识别缺失内容并制定获取与使用策略,上下文锚定则从推理、搜索、记忆和反馈中获取所需信息。为评估智能体图像生成能力,进一步构建Image Agent Bench(IA-Bench),涵盖计划、推理、搜索与记忆四大核心能力。在IA-Bench、Mindbench和WISE-Verified上的实验表明,Qwen-Image-Agent优于强基线,达到当前最佳性能。
原文摘要 · Abstract (English)
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowledge. We identify this challenge as the Context Gap: the mismatch between the user context and the sufficient generation context for T2I models. To bridge this gap, we propose Qwen-Image-Agent, a unified agentic framework that integrates plan, reason, search, memory and feedback in a context-centric manner. Qwen-Image-Agent treats user input as partial context and progressively constructs the generation context through Context-Aware Planning and Context Grounding. Specifically, Context-Aware Planning identifies missing context and plans how it should be acquired and used, while Context Grounding gathers this context from reason, search, memory, and feedback. To evaluate agentic image generation, we further introduce Image Agent Bench (IA-Bench), a benchmark covering four core image agent capabilities: Plan, Reason, Search, and Memory. Experiments on IA-Bench, Mindbench and WISE-Verified show that Qwen-Image-Agent outperforms strong baselines and achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。