DEI提示让医疗大模型虚构患者信息,可能误导诊疗决策。
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
- 在医疗问题后加一句DEI提示,模型会无中生有添加种族、性别等身份信息。
- 47个模型中,虚构率从0.7%飙升至33.1%,且99.8%的错误倾向导致答案偏差。
- 该现象由公平性内容引发,与提示长度无关,适用于所有测试模型。
临床AI指南建议通过提示引导语言模型关注多样性、公平性和包容性(DEI)。我们发现一种副作用:在医疗问题后添加一句简短的DEI提示,会使模型在未提及患者身份的情况下,主动引入种族、社会经济地位、性别等属性,实质上重写患者身份,称之为‘人口统计注入’。在47个模型、4个医学基准测试和37.6万条由验证过的模型裁判系统评分的回答中,单一DEI提示使注入率从0.7%上升至33.1%(提升47倍),且在全部47个模型中均显著发生,归因于公平性内容而非提示长度(比长度匹配对照组高出18倍;p=1.4×10⁻¹⁴)。多数新增内容为泛化人群陈述,不影响答案;但小部分将属性关联到具体患者或改变选项选择(占响应的0.25%-2.4%,99.8%倾向错误选项),导致模型推荐的治疗方案发生变化。提示措辞可使影响范围从14%扩大至56%。这表明DEI提示只是更普遍机制的一个例子:任何引导模型推理方式的指令都可能诱发其添加未经请求的细节,包括患者相关信息。被标记的输出被视为研究中的模型错误,而非临床建议。
原文摘要 · Abstract (English)
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。