arXiv:2510.21120cs.CV2025-10被引 1

通过生成反事实图像对,精准定位导致图片不安全的关键特征。

SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation

  • 用图像编辑模型只改影响安全性的特征,保持其他细节不变。
  • 构建超3000张反事实图像对,覆盖9类安全场景的细微差异。
  • 适合研究安全检测模型、训练轻量级防护模型的开发者使用。

什么特征让一张图片变得不安全?现有图像安全数据集标签粗略且模糊,无法识别具体触发风险的视觉元素。本文提出SafetyPairs框架,通过图像编辑模型生成仅在安全相关特征上不同的反事实图像对,实现安全标签的精准切换。该方法能系统性地分离出影响安全判断的微小变化,如一个侮辱性手势或符号。基于此,我们构建了一个新基准,包含超过3,020张图像,覆盖9种安全类别,可揭示视觉-语言模型在区分细微差异上的能力缺陷。此外,该方法作为数据增强策略,显著提升轻量级防护模型的训练效率。该基准已公开,是首个系统性研究图像安全细粒度差异的资源。

原文摘要 · Abstract (English)

What exactly makes a particular image unsafe? Systematically differentiating between benign and problematic images is a challenging problem, as subtle changes to an image, such as an insulting gesture or symbol, can drastically alter its safety implications. However, existing image safety datasets are coarse and ambiguous, offering only broad safety labels without isolating the specific features that drive these differences. We introduce SafetyPairs, a scalable framework for generating counterfactual pairs of images, that differ only in the features relevant to the given safety policy, thus flipping their safety label. By leveraging image editing models, we make targeted changes to images that alter their safety labels while leaving safety-irrelevant details unchanged. Using SafetyPairs, we construct a new safety benchmark, which serves as a powerful source of evaluation data that highlights weaknesses in vision-language models' abilities to distinguish between subtly different images. Beyond evaluation, we find our pipeline serves as an effective data augmentation strategy that improves the sample efficiency of training lightweight guard models. We release a benchmark containing over 3,020 SafetyPair images spanning a diverse taxonomy of 9 safety categories, providing the first systematic resource for studying fine-grained image safety distinctions.

图像安全反事实生成数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。