用自动注入瑕疵的方法生成像素级标注数据,提升图像生成质量。
Synthesizing Artifact Dataset for Pixel-level Detection
- 通过预设区域自动添加瑕疵,生成无需人工标注的像素级数据。
- 使检测器在真实标注数据上分别提升13.2%和3.7%的性能。
- 适合需要高质量图像生成与瑕疵检测的研究者使用。
artifact检测器能通过作为微调阶段的奖励模型,提升图像生成模型的输出保真度与美学质量。然而,训练此类检测器需昂贵的像素级人工标注,限制了其性能。直接使用弱检测器伪标注会引入噪声标签,导致效果不佳。为此,我们提出一种瑕疵注入管道,可自动在高质量合成图像的指定区域注入瑕疵,从而生成像素级标注而无需人工干预。该方法训练的检测器在人类标注数据上相较基线方法,对ConvNeXt提升13.2%,对Swin-T提升3.7%。本工作为可扩展的像素级瑕疵标注数据集奠定了基础,并融入世界知识以增强检测能力。
原文摘要 · Abstract (English)
Artifact detectors have been shown to enhance the performance of image-generative models by serving as reward models during fine-tuning. These detectors enable the generative model to improve overall output fidelity and aesthetics. However, training the artifact detector requires expensive pixel-level human annotations that specify the artifact regions. The lack of annotated data limits the performance of the artifact detector. A naive pseudo-labeling approach-training a weak detector and using it to annotate unlabeled images-suffers from noisy labels, resulting in poor performance. To address this, we propose an artifact corruption pipeline that automatically injects artifacts into clean, high-quality synthetic images on a predetermined region, thereby producing pixel-level annotations without manual labeling. The proposed method enables training of an artifact detector that achieves performance improvements of 13.2% for ConvNeXt and 3.7% for Swin-T, as verified on human-labeled data, compared to baseline approaches. This work represents an initial step toward scalable pixel-level artifact annotation datasets that integrate world knowledge into artifact detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。