用多智能体迭代推理,让文字生成图像更精准。
M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning
- 拆解提示词为可验证清单,分步修正约束条件
- 在OneIG-EN上达0.532得分,超越商用模型
- 无需重训练,可插件式提升任意文生图模型
生成模型在文本到图像合成中已实现高保真度,但在涉及多重约束的复杂组合提示上仍表现不佳。我们提出M3(多模态、多智能体、多轮视觉推理)框架,通过推理时迭代优化,系统性解决此类问题。M3在不重新训练的前提下,协调现成的基础模型形成稳健的多智能体循环:规划者将提示分解为可验证清单,检查者、优化者和编辑者智能逐项修正约束,验证者确保性能单调提升。应用于开源模型时,M3在挑战性OneIG-EN基准上表现优异,我们的Qwen-Image+M3超越Imagin4(0.515)与Seedream 3.0(0.530),达到0.532的顶尖水平。同时显著提升GenEval组合评估指标,硬化测试集上的空间推理能力翻倍。作为即插即用模块,M3为组合生成提供了无需昂贵重训练的新范式。
原文摘要 · Abstract (English)
Generative models have achieved impressive fidelity in text-to-image synthesis, yet struggle with complex compositional prompts involving multiple constraints. We introduce \textbf{M3 (Multi-Modal, Multi-Agent, Multi-Round)}, a training-free framework that systematically resolves these failures through iterative inference-time refinement. M3 orchestrates off-the-shelf foundation models in a robust multi-agent loop: a Planner decomposes prompts into verifiable checklists, while specialized Checker, Refiner, and Editor agents surgically correct constraints one at a time, with a Verifier ensuring monotonic improvement. Applied to open-source models, M3 achieves remarkable results on the challenging OneIG-EN benchmark, with our Qwen-Image+M3 surpassing commercial flagship systems including Imagen4 (0.515) and Seedream 3.0 (0.530), reaching state-of-the-art performance (0.532 overall). This demonstrates that intelligent multi-agent reasoning can elevate open-source models beyond proprietary alternatives. M3 also substantially improves GenEval compositional metrics, effectively doubling spatial reasoning performance on hardened test sets. As a plug-and-play module compatible with any pre-trained T2I model, M3 establishes a new paradigm for compositional generation without costly retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。