测试闭源大模型医学推理解释的可信度,发现其理由常不靠谱。
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
- 用三种扰动实验检测模型推理是否真实影响答案
- 多数情况下推理步骤不影响预测结果,易受外部提示干扰
- 适合关注AI医疗可信性的医生和研究者阅读
闭源大语言模型(如ChatGPT、Gemini)被越来越多用于医疗建议,但其生成的解释可能看似合理却与真实推理过程无关。这种偏差可能导致患者和临床医生误信误导性说明。本研究对三款主流闭源模型在医学推理中的可信度进行了系统性黑箱评估。采用三种基于扰动的探测方法:(1) 因果消融,检验陈述的思维链(CoT)是否真正影响预测;(2) 位置偏倚,考察模型是否因输入顺序而生成事后辩护;(3) 提示注入,测试模型对外部暗示的敏感度。结合小规模人工评估,分析医生对解释可信度的判断与普通人信任感之间的关联。结果表明,大多数情况下,思维链步骤并未因果性地驱动预测,模型极易采纳外部提示且无承认。而位置偏倚在此设置中影响较小。研究强调,在医疗领域评估大模型时,可信度应与准确性同等重要,以保障公众安全并实现临床安全部署。
原文摘要 · Abstract (English)
Closed-source large language models (LLMs), such as ChatGPT and Gemini, are increasingly consulted for medical advice, yet their explanations may appear plausible while failing to reflect the model's underlying reasoning process. This gap poses serious risks as patients and clinicians may trust coherent but misleading explanations. We conduct a systematic black-box evaluation of faithfulness in medical reasoning among three widely used closed-source LLMs. Our study consists of three perturbation-based probes: (1) causal ablation, testing whether stated chain-of-thought (CoT) reasoning causally influences predictions; (2) positional bias, examining whether models create post-hoc justifications for answers driven by input positioning; and (3) hint injection, testing susceptibility to external suggestions. We complement these quantitative probes with a small-scale human evaluation of model responses to patient-style medical queries to examine concordance between physician assessments of explanation faithfulness and layperson perceptions of trustworthiness. We find that CoT reasoning steps often do not causally drive predictions, and models readily incorporate external hints without acknowledgment. In contrast, positional biases showed minimal impact in this setting. These results underscore that faithfulness, not just accuracy, must be central in evaluating LLMs for medicine, to ensure both public protection and safe clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。