针对教育场景下视觉语言模型的隐性偏见,提出三层次审计框架并发现视觉输入会诱发文本防护失效
Edu-MMBias: A Three-Tier Multimodal Benchmark for Auditing Social Bias in Vision-Language Models under Educational Contexts

- 基于社会心理三维度构建认知-情感-行为偏见审计框架
- 发现模型对低社会地位群体有补偿性偏袒,同时存在健康与种族刻板印象
- 揭示视觉信息可绕过文本对齐机制,成为偏见复现的隐蔽通道
随着视觉语言模型(VLMs)在教育决策中日益重要,确保其公平性至关重要。然而现有以文本为中心的评估忽略了视觉模态,留下了潜在社会偏见的未监管通道。为此,我们提出Edu-MMBias,一个基于社会心理学态度三组件模型的系统性审计框架,从认知、情感和行为三个层级诊断偏见。通过包含自纠正机制与人机协同验证的专用生成流水线,我们合成抗污染的学生画像,对前沿VLMs进行全方位压力测试。大规模审计揭示出关键且反直觉的模式:模型对低社会地位叙事表现出补偿性偏见,同时深藏健康与种族刻板印象。关键发现是,视觉输入可作为安全后门,触发被文本对齐机制抑制的偏见重现,暴露出潜在认知与最终决策之间的系统性错位。相关成果已公开于https://anonymous.4open.science/r/EduMMBias-63B2。
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) become integral to educational decision-making, ensuring their fairness is paramount. However, current text-centric evaluations neglect the visual modality, leaving an unregulated channel for latent social biases. To bridge this gap, we present Edu-MMBias, a systematic auditing framework grounded in the tri-component model of attitudes from social psychology. This framework diagnoses bias across three hierarchical dimensions: cognitive, affective, and behavioral. Utilizing a specialized generative pipeline that incorporates a self-correct mechanism and human-in-the-loop verification, we synthesize contamination-resistant student profiles to conduct a holistic stress test on state-of-the-art VLMs. Our extensive audit reveals critical, counter-intuitive patterns: models exhibit a compensatory class bias favoring lower-status narratives while simultaneously harboring deep-seated health and racial stereotypes. Crucially, we find that visual inputs act as a safety backdoor, triggering a resurgence of biases that bypass text-based alignment safeguards and revealing a systematic misalignment between latent cognition and final decision-making. The contributions of this paper are available at: https://anonymous.4open.science/r/EduMMBias-63B2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。