通过多元视角对抗性生成,显著降低视觉语言模型的社交偏见。
Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

- 在视觉空间构建多群体反事实表征,打破刻板印象
- 解码时融合不同群体的高置信度输出,使结果更均衡
- 在职业、描述等场景中偏见降低最高达47.97%,适合关注公平性的研究者
大型视觉语言模型在多项任务中表现卓越,但常继承训练数据中的社会偏见,导致对不同社会群体肖像处理时产生偏差。现有去偏方法依赖单一刻板视角比较词元概率,难以覆盖社会观点多样性。受社会科学研究中多样性促进公平的启发,我们提出反事实集成解码(CED)框架:在视觉表示空间中构建多群体反事实视角,并在解码阶段融合这些视角。CED首先通过识别与各社会群体相关的语义方向,在视觉空间生成反事实表征,提供多样化视角以打破刻板叙事;解码时,定位不同视角间差异最大的解码层,使用不确定性感知权重集成其词元分布,优先保留来自不同群体的高置信度输出,从而生成更平衡的概率分布,引导更公平的生成结果。在三个社会偏见评估基准上的实验表明,相比领先基线, ool 在职业、描述和人物特质等场景下偏见减少最高达47.97%。此外,该方法对原模型核心能力影响极小,保持了良好性能。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups. Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives. Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior. CED first performs counterfactual steering in the visual space by identifying semantic directions associated with each social group and generating counterfactual representations along these directions, thereby offering diverse perspectives that disrupt stereotypical narratives. During decoding, CED locates the decoder layer exhibiting the greatest divergence among these perspectives and ensembles their token distributions using uncertainty-aware weights, prioritizing high-confidence tokens from different groups to yield a more balanced probability distribution that guides fairer generation. Extensive experiments on three social bias evaluation benchmarks demonstrate that \tool achieves substantial improvements over leading baselines, reducing bias by up to 47.97% across scenarios involving occupations, descriptors, and persona traits. Moreover, CED also preserves the core capabilities of the original model with minimal degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。