LLM在医疗文本中易被烟酒提及误导,误判药物使用情况。
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
- 通过提示工程和思维链减少模型对烟酒线索的依赖。
- 烟酒提及导致无药物史患者被误判为有用药史,错误率超20%。
- 发现模型对不同性别的判断存在明显偏差,适合医疗AI可信性研究者。
从临床文本中提取社会健康决定因素(SDOH)对下游医疗分析至关重要。尽管大语言模型(LLMs)表现良好,但可能依赖表面线索导致虚假预测。基于SHAC数据集中MIMIC部分,以药物状态提取为例,我们发现提及酒精或吸烟会错误诱导模型判断存在当前或既往药物使用,即使实际并无用药史;同时揭示了模型性能在性别上的显著差异。我们进一步评估了提示工程与思维链推理等缓解策略,有效降低此类误判,为提升医疗领域LLM的可靠性提供实证洞察。
原文摘要 · Abstract (English)
Social determinants of health (SDOH) extraction from clinical text is critical for downstream healthcare analytics. Although large language models (LLMs) have shown promise, they may rely on superficial cues leading to spurious predictions. Using the MIMIC portion of the SHAC (Social History Annotation Corpus) dataset and focusing on drug status extraction as a case study, we demonstrate that mentions of alcohol or smoking can falsely induce models to predict current/past drug use where none is present, while also uncovering concerning gender disparities in model performance. We further evaluate mitigation strategies - such as prompt engineering and chain-of-thought reasoning - to reduce these false positives, providing insights into enhancing LLM reliability in health domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。