arXiv:2604.17316cs.CLcs.AI2026-04ACL被引 2

身份标记让大模型医疗问答更不准且自信心失控

Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA

  • 用社会身份标签测试模型在2364个医学问题上的表现
  • 同性恋标签导致准确率下降,交叉身份危害非线性叠加
  • 错误不只影响答案,还扭曲了模型的自信信号

大型语言模型(LLMs)在临床安全部署中不仅需高准确率,还需可靠的不确定性校准,以确保模型在不确定时能合理依赖医生判断。本文研究患者社会属性(尤其是性取向与宗教信仰)如何扭曲模型的准确性与置信度信号。我们在2,364个医学问题及其反事实变体上评估了九种通用与生物医学领域大模型,发现身份标记引发“校准危机”:‘同性恋’标记持续降低性能,交叉身份产生非加性的独特伤害。此外,一项由临床医生验证的开放式生成案例研究证实,这些缺陷并非多选题格式所致。结果表明,社会身份线索不仅改变预测结果,更影响置信度信号的可靠性,对公平医疗与基于置信度的临床工作流构成重大风险。

原文摘要 · Abstract (English)

Safe clinical deployment of Large Language Models (LLMs) requires not only high accuracy but also robust uncertainty calibration to ensure models defer to clinicians when appropriate. Our paper investigates how social descriptors of a patient (specifically sexual orientation and religious affiliation) distort these uncertainty signals and model accuracy. Evaluating nine general-purpose and biomedical LLMs on 2,364 medical questions and their counterfactual variants, we demonstrate that identity markers cause a "calibration crisis". "Homosexual" markers consistently trigger performance drops, and intersectional identities produce idiosyncratic, non-additive harms to calibration. Moreover, a clinician-validated case study in an open-ended generation setting confirms that these failures are not an artifact of the multiple-choice format. Our results demonstrate that the presence of social identity cues does not merely shift predictions; it affects the reliability of confidence signals, posing a significant risk to equitable care and safe deployment in confidence-based clinical workflows.

大模型安全医疗AI偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。