测试大模型预测早产儿视网膜病变风险的能力与情感偏见。
Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- 构建中文数据集CROP,设计三种提示策略评估模型表现。
- 仅靠内在知识时模型效果差,引入外部知识后准确率显著提升。
- 积极情绪提示能缓解模型偏见,适合医疗AI公平性研究。
尽管大语言模型在多个领域取得显著进展,其在早产儿视网膜病变(ROP)风险预测方面的能力仍鲜有研究。为此,我们构建了首个中文基准数据集CROP,包含993条入院记录,标注为低、中、高风险三类。提出Affective-ROPTester框架,采用三种提示策略:指令式、思维链(CoT)和上下文学习(ICL),系统评估模型预测能力与情感偏见。指令式考察模型内在知识与偏见,而CoT与ICL利用外部医学知识提升准确性。关键创新在于在提示中引入情感元素,探究不同情绪语境对预测结果及偏见模式的影响。实证结果显示:仅依赖内在知识时,模型预测效果有限;引入结构化外部输入后性能显著提升;模型普遍存在对中高风险病例的高估倾向,且正向情绪提示相比负向情绪有助于降低预测偏差。这些发现强调了情感敏感型提示工程对提升临床诊断可靠性的重要性,并验证了Affective-ROPTester在评估与缓解医疗语言模型情感偏见方面的价值。
原文摘要 · Abstract (English)
Despite the remarkable progress of large language models (LLMs) across various domains, their capacity to predict retinopathy of prematurity (ROP) risk remains largely unexplored. To address this gap, we introduce a novel Chinese benchmark dataset, termed CROP, comprising 993 admission records annotated with low, medium, and high-risk labels. To systematically examine the predictive capabilities and affective biases of LLMs in ROP risk stratification, we propose Affective-ROPTester, an automated evaluation framework incorporating three prompting strategies: Instruction-based, Chain-of-Thought (CoT), and In-Context Learning (ICL). The Instruction scheme assesses LLMs' intrinsic knowledge and associated biases, whereas the CoT and ICL schemes leverage external medical knowledge to enhance predictive accuracy. Crucially, we integrate emotional elements at the prompt level to investigate how different affective framings influence the model's ability to predict ROP and its bias patterns. Empirical results derived from the CROP dataset yield two principal observations. First, LLMs demonstrate limited efficacy in ROP risk prediction when operating solely on intrinsic knowledge, yet exhibit marked performance gains when augmented with structured external inputs. Second, affective biases are evident in the model outputs, with a consistent inclination toward overestimating medium- and high-risk cases. Third, compared to negative emotions, positive emotional framing contributes to mitigating predictive bias in model outputs. These findings highlight the critical role of affect-sensitive prompt engineering in enhancing diagnostic reliability and emphasize the utility of Affective-ROPTester as a framework for evaluating and mitigating affective bias in clinical language modeling systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。