arXiv:2609.06058cs.CVcs.CL2026-09

权威提示让视觉语言模型偏离判断,即使被要求无视。

GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them

论文配图:GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them
图 1 · 摘自论文原文
  • 设计冲突场景测试模型在权威提示下的独立判断能力。
  • 官方答案使错误判断率上升19.5个百分点,远超其他提示。
  • 三款主流模型均受权威影响,适合评估AI可信度的研究者参考。

随着视觉语言模型(VLMs)在关键现实场景中应用日益广泛,其必须独立评估证据而非盲目服从人类权威。我们提出GradeTrap,一种受控评估方法,将学生答案(引发迎合性认同)与教师、同学或官方答案键(引发权威服从)置于直接冲突中。模型在明确指令下需独立作答并忽略所有学生回答、反馈和评分标记。在60个合成的真实世界权衡场景中测试,先通过5个中性试次建立稳定模型偏好基准,再重复三轮六类实验提示(含对照)。在Gemini 3.5 Flash-Lite、GPT-5.6 Luna和Claude Haiku 4.5共45项交集任务上,通用第二答案对照组导致5.4%的错误选择。相较于该对照,合并项分析显示:无显著同伴评审效应,教师评审效应为6.9分点,官方答案效应达19.5分点。相比之下,仅展示冲突学生答案相比仅展示学生参考答案,错误率从2.2%升至5.2%。尽管存在明确忽略指令及对立学生答案,官方答案来源仍显著改变判断,且效应大小在三模型间存在差异。

原文摘要 · Abstract (English)

As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct conflict: a student answer, which should attract sycophantic agreement, and a conflicting answer attributed to a peer, teacher, or official answer key, which should attract authority-based deference. Models produce free-form answers while being explicitly instructed to solve independently and ignore all student answers, feedback, and grading marks. We test the models on 60 synthetic real-world trade-off scenarios. Five neutral trials establish a stable model-relative preference, followed by three repetitions of six experimental cues including controls. On the 45-item common intersection across Gemini 3.5 Flash-Lite, GPT-5.6 Luna, and Claude Haiku 4.5, a generic second-answer control yields 5.4% conflicting-answer selection. Relative to that control, pooled within-item changes show no reliable peer-review effect, a 6.9-point teacher-review effect, and a 19.5-point official-key effect. In contrast, a displayed conflicting student answer alone compared to a displayed student reference answer alone only raises selection from 2.2% to 5.2%. Official-key provenance therefore redirects judgements more than a student answer or the generic second-answer control, despite an explicit ignore instruction and an opposing student answer given along with the official key. Effects vary in magnitude across the three models.

视觉语言模型权威偏差AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。