通过信息流分析揭示视觉语言模型的社会偏见根源
Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation
- 结合信息流分析与多轮对话,追踪模型推理中关键图像标记的贡献度
- 发现不同人群图像在模型内部的信息利用存在系统性差异
- 适用于研究模型偏见机制的AI安全与可解释性研究人员
大型视觉语言模型(LVLMs)在多模态任务中表现卓越,但亦存在显著社会偏见。这些偏见常表现为中性概念与敏感人类属性之间的无意关联,导致不同人口群体间模型行为差异。现有研究多聚焦于偏见的检测与量化,却难以揭示模型内部机制。为此,本文提出一种结合信息流分析与多轮对话评估的解释框架,旨在从内部信息利用不均衡的角度理解社会偏见成因。首先,通过信息流分析识别中性问题推理过程中高贡献的图像标记;随后设计多轮对话机制,评估这些关键标记是否编码敏感信息。大量实验表明,LVLM在处理不同人口群体图像时存在系统性信息使用差异,说明社会偏见深植于模型内部推理动态。此外,从文本模态补充验证:模型语义表示已呈现偏倚邻近模式,提供跨模态偏见形成解释。
原文摘要 · Abstract (English)
Large Vision Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet they also exhibit notable social biases. These biases often manifest as unintended associations between neutral concepts and sensitive human attributes, leading to disparate model behaviors across demographic groups. While existing studies primarily focus on detecting and quantifying such biases, they offer limited insight into the underlying mechanisms within the models. To address this gap, we propose an explanatory framework that combines information flow analysis with multi-round dialogue evaluation, aiming to understand the origin of social bias from the perspective of imbalanced internal information utilization. Specifically, we first identify high-contribution image tokens involved in the model's reasoning process for neutral questions via information flow analysis. Then, we design a multi-turn dialogue mechanism to evaluate the extent to which these key tokens encode sensitive information. Extensive experiments reveal that LVLMs exhibit systematic disparities in information usage when processing images of different demographic groups, suggesting that social bias is deeply rooted in the model's internal reasoning dynamics. Furthermore, we complement our findings from a textual modality perspective, showing that the model's semantic representations already display biased proximity patterns, thereby offering a cross-modal explanation of bias formation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。