用简化模型解析大模型在医疗预测中的隐藏知识,发现其可能携带错误偏见。
Surrogate modeling for interpreting black-box LLMs in medical predictions
- 通过大量模拟提示构建输入输出对,逼近大模型的隐含知识空间。
- 实验证明大模型会保留与医学常识相悖的关联和已被否定的种族假设。
- 适合关注AI医疗安全、可解释性及偏见检测的研究者使用。
大型语言模型(LLMs)在海量数据上训练,其参数中编码了丰富的现实世界知识,但其黑箱特性使这些知识的机制与范围难以理解。本文提出一种代理建模框架,可定量解析LLM编码的知识。针对基于领域知识的特定假设,该框架通过在广泛模拟场景下进行大量提示,利用可观测的输入-输出对逼近模型的潜在知识空间。在医疗预测的原型实验中,我们展示了该框架揭示大模型对各输入变量与输出关系“感知程度”的有效性。尤其值得注意的是,面对大模型可能延续训练数据中的错误信息与社会偏见的担忧,实验结果定量揭示了与公认医学知识矛盾的关联,以及科学上已被驳斥的种族假设在模型知识中的持续存在。该框架可作为预警信号,支持大模型在医疗等关键领域的安全可靠应用。
原文摘要 · Abstract (English)
Large language models (LLMs), trained on vast datasets, encode extensive real-world knowledge within their parameters, yet their black-box nature obscures the mechanisms and extent of this encoding. Surrogate modeling, which uses simplified models to approximate complex systems, can offer a path toward better interpretability of black-box models. We propose a surrogate modeling framework that quantitatively explains LLM-encoded knowledge. For a specific hypothesis derived from domain knowledge, this framework approximates the latent LLM knowledge space using observable elements (input-output pairs) through extensive prompting across a comprehensive range of simulated scenarios. Through proof-of-concept experiments in medical predictions, we demonstrate our framework's effectiveness in revealing the extent to which LLMs "perceive" each input variable in relation to the output. Particularly, given concerns that LLMs may perpetuate inaccuracies and societal biases embedded in their training data, our experiments using this framework quantitatively revealed both associations that contradict established medical knowledge and the persistence of scientifically refuted racial assumptions within LLM-encoded knowledge. By disclosing these issues, our framework can act as a red-flag indicator to support the safe and reliable application of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。