arXiv:2512.19735cs.LG2025-12

用历史病例引导大模型,提升重症预测公平性与准确性

Improving Fairness of Large Language Model-Based ICU Mortality Prediction via Case-Based Prompting

  • 基于相似病例配对纠正模型偏见,无需重新训练
  • 在MIMIC-IV数据集上AUC提升至0.873,偏见减少超90%
  • 适合关注临床决策公平性的医疗AI研究者

准确预测重症患者死亡风险对临床决策至关重要。尽管大语言模型在结构化医疗任务中展现出强大潜力,但其输出可能因性别、年龄、种族等人口属性产生偏差,限制了在公平性敏感的临床场景中的可靠性。现有去偏方法常导致预测性能下降,难以兼顾公平与准确。本文系统分析了基于大模型的重症死亡预测中的公平性问题,提出一种无需训练的临床自适应提示框架——案例提示(CAP),通过整合现有去偏策略,并引入相似历史误判案例及其正确结果来引导模型修正偏见推理。在MIMIC-IV数据集上的实验表明,AUROC从0.806提升至0.873,AUPRC从0.497增至0.694;不同人口群体间的预测差异显著降低,性别及白人-黑人比较中降幅超过90%。特征依赖分析显示各群体注意力模式高度一致,相似度高于0.98。结果表明,通过精心设计的提示可同时优化大模型临床预测的公平性与性能,为构建可靠且公平的临床辅助系统提供可行范式。

原文摘要 · Abstract (English)

Accurately predicting mortality risk in intensive care unit (ICU) patients is essential for clinical decision-making. Although large language models (LLMs) show strong potential in structured medical prediction tasks, their outputs may exhibit biases related to demographic attributes such as sex, age, and race, limiting their reliability in fairness-critical clinical settings. Existing debiasing methods often degrade predictive performance, making it difficult to balance fairness and accuracy. In this study, we systematically analyze fairness issues in LLM-based ICU mortality prediction and propose a clinically adaptive prompting framework that improves both performance and fairness without model retraining. We first design a multi-dimensional bias assessment scheme to identify subgroup disparities. Based on this, we introduce CAse Prompting (CAP), a training-free framework that integrates existing debiasing strategies and further guides models using similar historical misprediction cases paired with correct outcomes to correct biased reasoning. We evaluate CAP on the MIMIC-IV dataset. Results show that AUROC improves from 0.806 to 0.873 and AUPRC from 0.497 to 0.694. Meanwhile, prediction disparities are substantially reduced across demographic groups, with reductions exceeding 90% in sex and certain White-Black comparisons. Feature reliance analysis further reveals highly consistent attention patterns across groups, with similarity above 0.98. These findings demonstrate that fairness and performance in LLM-based clinical prediction can be jointly optimized through carefully designed prompting, offering a practical paradigm for developing reliable and equitable clinical decision-support systems.

大模型医疗公平性重症预测提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。