LLM在药物安全预测中会受用户背景影响,可能对弱势群体误判更高风险。
Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions
- 用真实报告数据测试模型对不同社会背景用户的药物副作用预测差异。
- 低教育、住房不稳群体被系统错误判定为更高风险,偏差显著且普遍存在。
- 揭示显性和隐性两种偏见模式,提醒医疗场景需警惕模型公平性问题。
大型语言模型(LLMs)在生物医学领域应用日益广泛,但其在药物安全预测中的可靠性仍待深入研究。本文探究了LLMs是否会在临床无关的背景下(如教育、婚姻状况、保险类型等)纳入社会人口学信息进行不良事件(AE)预测。基于美国食品药品管理局不良事件报告系统(FAERS)的结构化数据,结合基于角色的评估框架,我们评估了两款先进模型:ChatGPT-4o与Bio-Medical-Llama-3.8B。测试涵盖教育、婚姻状况、就业、保险、语言、住房稳定性及宗教等多维度人物画像,并考察全科医生、专科医生与患者三种用户角色的预测表现。结果发现,预测准确率存在系统性差异:低教育水平、住房不稳定的群体被频繁赋予更高的不良事件发生概率,而高学历、私有保险者则被低估。进一步分析识别出两种偏见模式:显性偏见——推理过程中直接引用人物属性;隐性偏见——预测不一致但未提及人物特征。这些发现暴露了将LLMs用于药物流行病学监测的重大风险,亟需建立公平性评估机制与缓解策略,以确保其在临床部署前的安全可靠。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied in biomedical domains, yet their reliability in drug-safety prediction remains underexplored. In this work, we investigate whether LLMs incorporate socio-demographic information into adverse event (AE) predictions, despite such attributes being clinically irrelevant. Using structured data from the United States Food and Drug Administration Adverse Event Reporting System (FAERS) and a persona-based evaluation framework, we assess two state-of-the-art models, ChatGPT-4o and Bio-Medical-Llama-3.8B, across diverse personas defined by education, marital status, employment, insurance, language, housing stability, and religion. We further evaluate performance across three user roles (general practitioner, specialist, patient) to reflect real-world deployment scenarios where commercial systems often differentiate access by user type. Our results reveal systematic disparities in AE prediction accuracy. Disadvantaged groups (e.g., low education, unstable housing) were frequently assigned higher predicted AE likelihoods than more privileged groups (e.g., postgraduate-educated, privately insured). Beyond outcome disparities, we identify two distinct modes of bias: explicit bias, where incorrect predictions directly reference persona attributes in reasoning traces, and implicit bias, where predictions are inconsistent, yet personas are not explicitly mentioned. These findings expose critical risks in applying LLMs to pharmacovigilance and highlight the urgent need for fairness-aware evaluation protocols and mitigation strategies before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。