arXiv:2501.12382cs.CV2025-01ICCV被引 7

定位图像生成中的瑕疵位置,帮助模型修复问题

DiffDoctor: Diagnosing Image Diffusion Models Before Treating

  • 构建百万级缺陷图像数据集,训练精准的瑕疵检测器
  • 在文本到图像生成中减少90%以上可见瑕疵
  • 适合需要高质量图像输出的研究者与开发者

尽管近期取得进展,图像扩散模型仍会产生伪影。现有方法通常依赖质量评估系统或人工评分整体图像进行优化。本文认为,解决问题应始于准确定位——模型不仅需感知缺陷存在,还需明确其具体位置。为此,我们提出 DiffDoctor:一种两阶段流程,以诊断并改善图像生成质量。第一阶段构建鲁棒的伪影检测器,收集超100万张合成缺陷图像,并设计类平衡策略的高效人机标注流程。第二阶段利用该检测器提供像素级反馈,优化扩散模型。在文本到图像扩散模型上的大量实验表明,所提检测器有效,且“先诊断后治疗”的设计合理可靠。

原文摘要 · Abstract (English)

In spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving starts with identification, yielding the request that the model should be aware of not just the presence of defects in an image, but their specific locations. Motivated by this, we propose DiffDoctor, a two-stage pipeline to assist image diffusion models in generating fewer artifacts. Concretely, the first stage targets developing a robust artifact detector, for which we collect a dataset of over 1M flawed synthesized images and set up an efficient human-in-the-loop annotation process, incorporating a carefully designed class-balance strategy. The learned artifact detector is then involved in the second stage to optimize the diffusion model by providing pixel-level feedback. Extensive experiments on text-to-image diffusion models demonstrate the effectiveness of our artifact detector as well as the soundness of our diagnose-then-treat design.

图像生成扩散模型瑕疵检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。