arXiv:2512.09081cs.CV2025-12被引 4

让文生图模型更懂复杂描述关系,生成更精准的图像。

AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models

  • 用大模型自动生成带细粒度差异的图文数据集
  • 在T2I-CompBench上达到当前最佳表现
  • 不损失画质,还能提升文本渲染等泛化能力

文生图模型虽已具备出色视觉质量,但在组合性方面仍存短板——难以准确捕捉提示词中的对象关系、属性绑定及细粒度细节。核心问题在于模型未显式学习区分语义相近但结构不同的提示与图像。为此,我们提出AgentComp框架,利用具备图像生成、编辑和视觉问答能力的大语言模型,自主构建组合性数据集,并通过代理偏好优化方法微调文生图模型,增强其对细微组合差异的区分能力。实验表明,AgentComp在T2I-CompBench等组合性基准上达到当前最优性能,且不牺牲图像质量,甚至在未显式训练的文本渲染任务中也表现出良好泛化能力。

原文摘要 · Abstract (English)

Text-to-image generative models have achieved remarkable visual quality but still struggle with compositionality$-$accurately capturing object relationships, attribute bindings, and fine-grained details in prompts. A key limitation is that models are not explicitly trained to differentiate between compositionally similar prompts and images, resulting in outputs that are close to the intended description yet deviate in fine-grained details. To address this, we propose AgentComp, a framework that explicitly trains models to better differentiate such compositional variations and enhance their reasoning ability. AgentComp leverages the reasoning and tool-use capabilities of large language models equipped with image generation, editing, and VQA tools to autonomously construct compositional datasets. Using these datasets, we apply an agentic preference optimization method to fine-tune text-to-image models, enabling them to better distinguish between compositionally similar samples and resulting in overall stronger compositional generation ability. AgentComp achieves state-of-the-art results on compositionality benchmarks such as T2I-CompBench, without compromising image quality$-$a common drawback in prior approaches$-$and even generalizes to other capabilities not explicitly trained for, such as text rendering.

文生图组合性大模型生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。