arXiv:2605.26353cs.CVcs.AI2026-05

用生成模型补足罕见场景数据,提升模型泛化能力

Personalized Generative Models for Contextual Debiasing

论文配图:Personalized Generative Models for Contextual Debiasing
图 1 · 摘自论文原文
  • 个性化扩散模型生成稀有场景图像,保持原数据视觉细节
  • 在复杂场景数据集上分类准确率显著优于基线方法
  • 适合需要提升模型鲁棒性的计算机视觉研究者

现实世界中不同视觉模式出现频率不一:例如沙滩球更常出现在沙地上而非道路上。这种统计偏差反映在视觉数据集中,导致训练模型更容易识别常见场景中的物体。然而,在道路等罕见场景中识别沙滩球可能更具实际意义。为缓解这一偏差,我们探索通过生成低频上下文图像作为训练增强。核心挑战在于生成多样且语义合理的罕见场景图像,同时保持与原始数据分布一致。为此提出DecoupleGen方法,通过个性化文本到图像扩散模型,实现对罕见上下文的连贯合成,保留原始视觉细节。生成图像包含有意义的内容,并通过验证约束确保数据相关性。在复杂场景数据集上的对象分类与识别任务中评估,实验显示性能持续优于现有方法,分析揭示了提升的关键因素。

原文摘要 · Abstract (English)

Different visual patterns appear with different frequencies in the world: e.g., beach balls appear on sand more often than they do on a road. These statistics are reflected in vision datasets, and as a result trained models more easily recognize objects in common scenarios. However, recognizing a beach ball on a road may arguably be even more important than recognizing it on sand. We study how to mitigate this discrepancy. Since collecting uncommon images in the real world may be difficult, we explore whether generating images with less frequent contexts can serve as effective training augmentation. A key challenge is guiding generations to remain close to the original dataset distribution while creating diverse images with uncommon contexts. We introduce Decoupling Contextual Patterns with Generations (DecoupleGen), a method that personalizes text-to-image diffusion models to facilitate coherent synthesis of images with rare contexts while preserving original visual details. The generated images contain semantically meaningful content and remain visually aligned with the original datasets. We further apply verification constraints to ensure relevance of the augmented data. We evaluate our approach on object classification and recognition tasks on complex scene datasets. Our experiments demonstrate consistent improvements over previous approaches, and our analyses identify factors underlying these improvements.

生成模型去偏扩散模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。