临床语言模型会因表达风格不同而误诊,研究提出新方法消除这种偏见。
Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

- 构建1000个改写病例,保持临床事实一致但表达风格不同,测试模型偏差。
- 7个模型均出现显著诊断差异,偏差值达0.064至0.151,影响诊疗一致性。
- 提出NarrativeShield框架,可将偏差降至近零,适合医疗AI可靠性研究者。
用于临床诊断推理的大语言模型对社会语言风格敏感,而非仅依赖临床内容。我们称此为叙事锚定:相同临床事实以不同语言风格表达时,模型输出诊断结果产生偏离。不同于以往针对种族、收入等显性身份标签的偏见研究,本工作仅改变语言风格,不引入任何人口统计学标记。我们构建了1000个美国医学执照考试(USMLE)临床案例,每个案例被重写为三种社会语言风格迥异的人物表述,在独立审计下确保事实不变,由另一模型验证未见生成提示。在涵盖三种架构和多个规模的七种语言模型中,直接提示下叙事锚定效应普遍存在,偏差值(Narrative Anchoring Gap)在0.064至0.151之间。链式思维推理与显式去偏指令仅部分缓解偏见,且常伴随准确率下降。我们提出NarrativeShield三代理器管道,在诊断前结构化提取并验证临床事实,使偏差降至接近零(-0.004至0.037),并实现所有模型中最低的严重决策不稳定性(DSS < 0.8)。在非指令微调的基础模型上进行压力测试发现,去偏干预能否生效取决于零样本指令遵循能力,而非提示内容本身。数据集已发布,经人工验证事实一致性,可作为注册式临床偏见研究的独立资源。
原文摘要 · Abstract (English)
Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。