提出结构化缺陷定位方法,精准诊断图像生成中的问题位置、类型、原因和重要性。
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

- 将缺陷视为位置、类型、原因、重要性的四元组,实现结构化预测
- 构建30K图像数据集并设计专用评估协议,支持实例级缺陷分析
- 通过视觉语言模型检测缺陷,并生成空间奖励提升图像生成质量
尽管文本到图像(T2I)模型生成的图像日益逼真,但仍存在局部、细微且结构复杂的错误。诊断这些错误需要实例级反馈,回答缺陷发生的位置、类型、成因及其对整体图像质量的重要性。现有密集反馈方法虽超越标量监督,但以热图为中心的表示仍将其视为像素场回归,难以定位可变数量的缺陷,也难以将语义原因与具体错误关联。为此,本文提出结构化缺陷定位(SDG),将T2I诊断建模为结构化集合预测,每个缺陷表示为(位置、类型、原因、重要性)四元组。为使该形式可训练可度量,我们构建了包含30,000张图像的SDG-30K数据集,覆盖四种现代T2I生成器,并提出专门的评估协议SDG-Eval。在此基础上,我们进一步提出诊断对齐框架:视觉语言模型(VLM)作为缺陷检测器,BoxFlow-GRPO将预测的缺陷集合转换为基于框的、加权的空间奖励,用于扩散模型对齐。大量实验表明,所提SDG检测器在结构化缺陷定位上优于领先专有VLM,且由SDG引导的奖励持续提升T2I对齐效果,支持局部图像优化。结果确立了SDG作为诊断、评估和增强现代生成模型的统一实例级接口。
原文摘要 · Abstract (English)
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality. While recent dense-feedback methods move beyond scalar supervision, their heatmap-centric representations still formulate diagnosis as pixel-field regression, making it difficult to localize variable-cardinality defects and bind semantic reasons to individual failures. To address this representation bottleneck, we propose Structured Defect Grounding (SDG), which casts T2I diagnosis as structured set prediction by modeling each defect as a (location, type, reason, importance) tuple. To make this formulation trainable and measurable, we introduce SDG-30K, a 30K-image dataset with box-grounded annotations across four modern T2I generators, together with a dedicated evaluation protocol, SDG-Eval. Building on this structured representation, we further present a diagnosis-to-alignment framework in which a Vision-Language Model (VLM) serves as the SDG detector, and BoxFlow-GRPO converts predicted defect sets into box-derived, importance-weighted spatial rewards for diffusion model alignment. Extensive experiments show that our SDG detector outperforms leading proprietary VLMs on structured defect grounding, while SDG-guided rewards consistently improve T2I alignment and support localized image refinement. These results establish SDG as a unified, instance-level interface for diagnosing, evaluating, and enhancing modern generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。