arXiv:2503.10486cs.CLcs.AI2025-03被引 11

对比两款大模型在慢病诊断中的表现,发现各有专长且存在明显短板。

LLMs in Disease Diagnosis: A Comparative Study of DeepSeek-R1 and O3 Mini Across Chronic Health Conditions

  • 用症状数据对比DeepSeek R1与O3 Mini的诊断准确率
  • DeepSeek R1整体准确率82%,在精神、神经、肿瘤领域达100%
  • 模型对呼吸系统疾病识别差,且自信预测比例差异显著

大型语言模型(LLMs)正在革新医疗诊断,提升疾病分类与临床决策能力。本研究评估了两款基于LLM的诊断工具DeepSeek R1和O3 Mini在结构化症状与诊断数据集上的表现,考察其在疾病级别和类别级别的预测准确率及置信度可靠性。DeepSeek R1在疾病级别准确率为76%,整体准确率为82%,优于O3 Mini的72%和75%。值得注意的是,DeepSeek R1在精神健康、神经系统疾病和肿瘤领域达到100%准确率,而O3 Mini在自身免疫性疾病分类中表现优异,准确率达100%。然而,两者在呼吸系统疾病诊断上表现不佳,分别为40%和20%。此外,信心评分分析显示,DeepSeek R1在92%的案例中给出高置信度预测,而O3 Mini为68%。研究还讨论了偏见、模型可解释性与数据隐私等伦理问题,以确保LLM在临床实践中的负责任应用。总体而言,本研究揭示了基于LLM的诊断系统的优劣,并为未来人工智能驱动的医疗优化提供了路线图。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are revolutionizing medical diagnostics by enhancing both disease classification and clinical decision-making. In this study, we evaluate the performance of two LLM- based diagnostic tools, DeepSeek R1 and O3 Mini, using a structured dataset of symptoms and diagnoses. We assessed their predictive accuracy at both the disease and category levels, as well as the reliability of their confidence scores. DeepSeek R1 achieved a disease-level accuracy of 76% and an overall accuracy of 82%, outperforming O3 Mini, which attained 72% and 75% respectively. Notably, DeepSeek R1 demonstrated exceptional performance in Mental Health, Neurological Disorders, and Oncology, where it reached 100% accuracy, while O3 Mini excelled in Autoimmune Disease classification with 100% accuracy. Both models, however, struggled with Respiratory Disease classification, recording accuracies of only 40% for DeepSeek R1 and 20% for O3 Mini. Additionally, the analysis of confidence scores revealed that DeepSeek R1 provided high-confidence predictions in 92% of cases, compared to 68% for O3 Mini. Ethical considerations regarding bias, model interpretability, and data privacy are also discussed to ensure the responsible integration of LLMs into clinical practice. Overall, our findings offer valuable insights into the strengths and limitations of LLM-based diagnostic systems and provide a roadmap for future enhancements in AI-driven healthcare.

大模型诊断慢病识别医疗AI模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。