用知识图谱和智能代理发现医疗大模型的隐蔽偏见
Toward Revealing Nuanced Biases in Medical LLMs
- 结合知识图谱与智能代理,生成更有效的测试问题
- 在三数据集六模型上,检测出比现有方法多30%以上的复杂偏见
- 适合医疗AI安全评估、模型审计人员使用
应用于医疗的大语言模型(LLM)普遍存在偏见和不公平现象。在临床决策中部署前,必须识别这些偏见以有效缓解并减少负面影响。本研究提出一种新框架,融合知识图谱(KG)与辅助性(代理式)LLM,系统揭示医疗LLM中的复杂偏见模式。该方法结合对抗性扰动(红队测试)技术,识别细微偏见,并采用定制化的多跳知识图谱表征,提升对目标LLM的系统评估能力。不仅生成更具针对性的红队测试问题,还更高效地利用这些问题揭示深层偏见。在三个数据集、六种模型和五类偏见上的综合实验表明,该框架相比其他常见方法,在揭示复杂偏见方面表现出显著更强的能力与可扩展性。
原文摘要 · Abstract (English)
Large language models (LLMs) used in medical applications are known to be prone to exhibiting biased and unfair patterns. Prior to deploying these in clinical decision-making, it is crucial to identify such bias patterns to enable effective mitigation and minimize negative impacts. In this study, we present a novel framework combining knowledge graphs (KGs) with auxiliary (agentic) LLMs to systematically reveal complex bias patterns in medical LLMs. The proposed approach integrates adversarial perturbation (red teaming) techniques to identify subtle bias patterns and adopts a customized multi-hop characterization of KGs to enhance the systematic evaluation of target LLMs. It aims not only to generate more effective red-teaming questions for bias evaluation but also to utilize those questions more effectively in revealing complex biases. Through a series of comprehensive experiments on three datasets, six LLMs, and five bias types, we demonstrate that our proposed framework exhibits a noticeably greater ability and scalability in revealing complex biased patterns of medical LLMs compared to other common approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。