LLM诊断类风湿性关节炎准确率95%,但推理错误率达68%。
Right Prediction, Wrong Reasoning: Uncovering LLM Misalignment in RA Disease Diagnosis
- 用多轮分析验证不同LLM在真实患者数据上的诊断表现。
- 模型预测准确率约95%,但专家评估其推理错误率达68%。
- 揭示高精度预测与错误推理间的严重不一致,警示临床应用风险。
大型语言模型(LLMs)在疾病早期筛查中展现出巨大潜力,可提升医疗可及性。本文基于真实患者数据研究了LLMs对类风湿性关节炎(RA)的诊断能力,将模型预测结果与医学专家诊断进行对比。结果显示,最佳模型对RA的预测准确率约为95%。然而,当医学专家评估模型生成的推理过程时,发现近68%的推理内容存在错误。这一发现揭示了模型在预测与解释之间存在显著脱节:尽管答案正确,但推理路径不可靠。该研究凸显了在临床决策中依赖模型解释的风险,强调需警惕高精度预测背后潜在的错误逻辑。
原文摘要 · Abstract (English)
Large language models (LLMs) offer a promising pre-screening tool, improving early disease detection and providing enhanced healthcare access for underprivileged communities. The early diagnosis of various diseases continues to be a significant challenge in healthcare, primarily due to the nonspecific nature of early symptoms, the shortage of expert medical practitioners, and the need for prolonged clinical evaluations, all of which can delay treatment and adversely affect patient outcomes. With impressive accuracy in prediction across a range of diseases, LLMs have the potential to revolutionize clinical pre-screening and decision-making for various medical conditions. In this work, we study the diagnostic capability of LLMs for Rheumatoid Arthritis (RA) with real world patients data. Patient data was collected alongside diagnoses from medical experts, and the performance of LLMs was evaluated in comparison to expert diagnoses for RA disease prediction. We notice an interesting pattern in disease diagnosis and find an unexpected \textit{misalignment between prediction and explanation}. We conduct a series of multi-round analyses using different LLM agents. The best-performing model accurately predicts rheumatoid arthritis (RA) diseases approximately 95\% of the time. However, when medical experts evaluated the reasoning generated by the model, they found that nearly 68\% of the reasoning was incorrect. This study highlights a clear misalignment between LLMs high prediction accuracy and its flawed reasoning, raising important questions about relying on LLM explanations in clinical settings. \textbf{LLMs provide incorrect reasoning to arrive at the correct answer for RA disease diagnosis.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。