arXiv:2602.23783cs.CV2026-02被引 10

用早期注意力图预测图像质量,提前淘汰低质生成结果。

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

  • 通过分析扩散模型初期的交叉注意力分布,建立质量预测机制。
  • 在多个模型和指标上相关性超0.7,分类准确率超0.9。
  • 适合需要高效迭代的生成场景,如提示词优化与强化学习训练。

文本到图像扩散模型缺乏高效的早期质量评估机制,导致提示词迭代、基于代理的生成等多生成场景中计算成本高昂。本文揭示了早期扩散交叉注意力分布与最终图像质量之间的强相关性。基于此,提出Diffusion Probe框架,利用内部交叉注意力图作为预测信号。设计轻量级预测器,将初始去噪步骤中提取的注意力统计特征映射为最终图像整体质量。该方法可在完整合成前准确预测多种评价指标下的图像质量。在多个T2I模型、不同去噪窗口、分辨率及质量度量下,相关性(PCC > 0.7)和分类性能(AUC-ROC > 0.9)均表现优异。其可靠性带来实际收益:支持提示词优化、种子选择与加速强化学习训练中的早筛决策,减少无效生成计算,提升最终输出质量。Diffusion Probe具有模型无关性、高效性和广泛适用性,为提升T2I生成效率提供了实用解决方案。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion models lack an efficient mechanism for early quality assessment, leading to costly trial-and-error in multi-generation scenarios such as prompt iteration, agent-based generation, and flow-grpo. We reveal a strong correlation between early diffusion cross-attention distributions and final image quality. Based on this finding, we introduce Diffusion Probe, a framework that leverages internal cross-attention maps as predictive signals. We design a lightweight predictor that maps statistical properties of early-stage cross-attention extracted from initial denoising steps to the final image's overall quality. This enables accurate forecasting of image quality across diverse evaluation metrics long before full synthesis is complete. We validate Diffusion Probe across a wide range of settings. On multiple T2I models, across early denoising windows, resolutions, and quality metrics, it achieves strong correlation (PCC > 0.7) and high classification performance (AUC-ROC > 0.9). Its reliability translates into practical gains. By enabling early quality-aware decisions in workflows such as prompt optimization, seed selection, and accelerated RL training, the probe supports more targeted sampling and avoids computation on low-potential generations. This reduces computational overhead while improving final output quality.Diffusion Probe is model-agnostic, efficient, and broadly applicable, offering a practical solution for improving T2I generation efficiency through early quality prediction.

图像生成扩散模型质量预测效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。