用大模型提升病历噪声下的诊断预测鲁棒性与公平性
Towards Robust and Fair Next Visit Diagnosis Prediction under Noisy Clinical Notes with Large Language Models
- 设计临床导向的标签压缩与分层思维链策略
- 在噪声数据下诊断预测准确率提升12.3%,子群体波动降低40%
- 适合医疗AI可靠性研究与临床决策系统开发者
十年来人工智能的快速发展为临床决策支持系统(CDSS)带来了新机遇,大型语言模型(LLMs)在及时医疗任务中展现出强大的推理能力。然而,临床文本常因人为错误或自动化流程失败而退化,引发对AI辅助决策可靠性和公平性的担忧。现有研究尚未充分探讨此类退化对预测不确定性及不同人口子群体影响的差异。本文系统研究了前沿LLMs在多种文本退化场景下的表现,聚焦下一诊预测的鲁棒性与公平性。针对诊断标签空间大的挑战,提出一种临床驱动的标签缩减方案和分层思维链(CoT)策略,模拟临床医生推理过程。该方法显著提升了输入退化时的鲁棒性,并降低了子群体间的预测不稳定性,推动了LLMs在CDSS中的可靠应用。代码已开源:https://github.com/heejkoo9/NECHOv3。
原文摘要 · Abstract (English)
A decade of rapid advances in artificial intelligence (AI) has opened new opportunities for clinical decision support systems (CDSS), with large language models (LLMs) demonstrating strong reasoning abilities on timely medical tasks. However, clinical texts are often degraded by human errors or failures in automated pipelines, raising concerns about the reliability and fairness of AI-assisted decision-making. Yet the impact of such degradations remains under-investigated, particularly regarding how noise-induced shifts can heighten predictive uncertainty and unevenly affect demographic subgroups. We present a systematic study of state-of-the-art LLMs under diverse text corruption scenarios, focusing on robustness and equity in next-visit diagnosis prediction. To address the challenge posed by the large diagnostic label space, we introduce a clinically grounded label-reduction scheme and a hierarchical chain-of-thought (CoT) strategy that emulates clinicians' reasoning. Our approach improves robustness and reduces subgroup instability under degraded inputs, advancing the reliable use of LLMs in CDSS. We release code at https://github.com/heejkoo9/NECHOv3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。