arXiv:2609.05540cs.CVcs.AI2026-09

测试视觉语言模型在无法判断自闭症时是否拒绝回答,发现部分表情会诱使模型错误推断。

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

论文配图:Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models
图 1 · 摘自论文原文
  • 构建带控因的合成人脸图像对,检验模型在无诊断依据时是否拒答。
  • 多数模型拒绝回答,但部分模型受表情影响,错误判定自闭症风险。
  • 改进提示词和界面设计可显著提升模型拒答率,保障医疗安全。

许多医疗判断需基于多维度临床评估,而非仅凭视觉外观。尽管如此,视觉语言模型(VLMs)常被用于解读医学相关图像,引发安全风险。自闭症(ASD)诊断依赖行为与发育证据,而非静态面部照片。我们评估VLMs在面对无法从图像中回答的配对图像问题时是否会主动拒答,并分析表情是否影响非拒答选择。提出PARITY(Paired Assessment with Reused Identity)数据集:一组身份控制、人口统计均衡的中性/表情人脸配对图像,以及中性-中性对照组。所有身份为合成生成,无实际ASD状态;因此任何非拒答的选择均视为有害误判。在主流VLM中,观察到明确的拒答优先与推测型模型分野;后者中特定表情显著增加错误判断概率。临床防护机制与单图呈现方式能显著提升拒答率,表明提示工程与界面设计具有可操作的缓解潜力。

原文摘要 · Abstract (English)

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such inferences are unsupported. ASD diagnosis requires behavioral and developmental evidence, not static facial photographs. We audit whether VLMs abstain from this unanswerable paired-image query, and whether expressions sway non-abstaining choices. We introduce PARITY (Paired Assessment with Reused Identity), a synthetic, demographically balanced set of identity-controlled neutral/expression portrait pairs with neutral-neutral controls. All identities are synthetic and have no ASD status; because the query is unanswerable from images, any non-abstaining selection is treated as a harmful attribution. Across contemporary VLMs, we find a clear split between refusal-first models and speculative models; in the latter, certain expressions disproportionately trigger harmful selections. Clinical guardrails and single-image framing substantially increase abstention, suggesting actionable mitigations in both prompting and interface design

视觉语言模型医疗安全模型拒答自闭症识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。