医学大模型会因患者叙述方式不同而改变答案,影响临床可信度。
SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

- 用角色化病历对比测试模型响应稳定性
- 70亿参数模型准确率56.33%,但仍有31.67%叙述敏感误差
- 评估临床决策支持需兼顾正确性与叙述一致性
医学大语言模型常以回答正确率评判,但这忽略了实际风险:模型可能知道正确答案,却因同一病例的不同患者叙述方式而改变回应。本文评估这种‘社会决定健康因素感知的叙事锚定偏差’。我们使用NarrativeShield SDoH MedQA这一反事实医疗问答数据集,其中每个病例以不同角色化叙述呈现,但答案保持不变。数据集从宽格式重构为按病例分组的角色行。评估Qwen2.5系列三个开源指令微调模型(1.5B、3B、7B),共使用300个临床案例,生成8,100条模型响应,涵盖三种提示条件。报告了角色级准确率、反事实一致性、正确一致性及叙事敏感性误差。Qwen2.5 7B在准确率上表现最佳(56.33%),正确一致性最佳(40.33%)。配对McNemar精确检验显示,7B在所有提示设置下均显著优于3B。然而,叙事敏感性误差仍存在,最低为31.67%。结果表明,可信临床决策支持应同时评估平均正确率和对医学等价叙述的稳定性。
原文摘要 · Abstract (English)
Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in a different patient voice. This paper evaluates that risk as SDoH aware narrative anchoring bias. We use NarrativeShield SDoH MedQA, a counterfactual medical question answering dataset in which each case appears in persona based narratives while the answer key remains fixed. The dataset is reshaped from wide format into case grouped persona rows. We evaluate three open source instruction tuned LLMs from the Qwen2.5 family: 1.5B, 3B, and 7B. The final experiment uses 300 clinical cases and produces 8,100 model responses across three prompting conditions. We report persona level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. Qwen2.5 7B achieves the best accuracy at 56.33 percent and the best correct consistency at 40.33 percent. Paired McNemar exact tests show significant accuracy gains for 7B over 3B in all prompt settings. Even so, narrative sensitivity remains, with the lowest error still at 31.67 percent. These results suggest that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。