arXiv:2412.09160cs.CV2024-12被引 1

通过局部化反事实生成,减少大模型中的社会偏见。

Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation

  • 仅修改与属性相关的局部区域,保持图像上下文不变。
  • 生成的反事实图像在视觉和语义上更真实,偏见降低显著。
  • 适合用于偏见分析与模型鲁棒性提升的研究者。

在网页抓取数据集上训练的基础模型会将社会偏见带入下游任务。尽管反事实生成可用于偏见分析,但现有方法因修改衣物、背景等上下文元素而引入伪影。我们提出一种局部化反事实生成方法,通过自动掩码和引导修复,将修改限制在特定属性相关区域,从而保留图像上下文。在Conceptual Captions数据集上生成性别反事实时,该方法在视觉与语义保真度上优于现有最优方案,且在非以人类为中心的任务中保持与仅使用真实数据训练模型相当的性能。用反事实数据微调的模型在多个指标上表现出可测量的偏见减少,包括性别分类差异下降及人物偏好得分均衡,同时维持ImageNet零样本性能。结果建立了可实现准确偏见剖析与有效缓解的平衡数据集框架。

原文摘要 · Abstract (English)

Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce artifacts by modifying contextual elements like clothing and background. We present a localized counterfactual generation method that preserves image context by constraining counterfactual modifications to specific attribute-relevant regions through automated masking and guided inpainting. When applied to the Conceptual Captions dataset for creating gender counterfactuals, our method results in higher visual and semantic fidelity than state-of-the-art alternatives, while maintaining the performance of models trained using only real data on non-human-centric tasks. Models fine-tuned with our counterfactuals demonstrate measurable bias reduction across multiple metrics, including a decrease in gender classification disparity and balanced person preference scores, while preserving ImageNet zero-shot performance. The results establish a framework for creating balanced datasets that enable both accurate bias profiling and effective mitigation.

偏见缓解反事实生成图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。