用视觉蓝图生成新图像,解决目标检测中罕见物体的偏差问题。
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
- 用表示分数诊断数据缺失,指导生成更均衡的图像布局。
- 生成图像中大物体和罕见物体分别提升4.4和3.6 mAP。
- 通过视觉蓝图与检测器协同生成,提升图像真实感与精度。
本文提出一种基于生成的去偏框架用于目标检测。以往去偏方法受限于样本表示多样性,而简单的生成增强往往保留原有偏差。我们分析发现,仅对稀有类别生成更多数据并非最优:其一,实例频率无法完整反映模型真实的数据需求;其二,当前布局到图像的合成缺乏保真度与控制力,难以生成高质量复杂场景。为此,我们引入表示分数(RS)以诊断超出频率的表征差距,并指导生成无偏布局。为保证合成质量,我们用精确的视觉蓝图替代模糊文本提示,并采用生成对齐策略,促进检测器与生成器之间的交互。实验表明,该方法显著缩小了低频物体组的性能差距,在大物体和罕见物体上分别比基线提升4.4和3.6 mAP;在生成图像的布局准确性上,相比先前的L2I合成模型提升15.9 mAP。
原文摘要 · Abstract (English)
This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generating more data for rare classes is suboptimal due to two core issues: i) instance frequency is an incomplete proxy for the true data needs of a model, and ii) current layout-to-image synthesis lacks the fidelity and control to generate high-quality, complex scenes. To overcome this, we introduce the representation score (RS) to diagnose representational gaps beyond mere frequency, guiding the creation of new, unbiased layouts. To ensure high-quality synthesis, we replace ambiguous text prompts with a precise visual blueprint and employ a generative alignment strategy, which fosters communication between the detector and generator. Our method significantly narrows the performance gap for underrepresented object groups, \eg, improving large/rare instances by 4.4/3.6 mAP over the baseline, and surpassing prior L2I synthesis models by 15.9 mAP for layout accuracy in generated images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。