arXiv:2512.17559cs.AI2025-12被引 1

用大模型打造可解释的问诊聊天机器人,提升诊断透明度与准确性。

Towards Explainable Conversational AI for Early Diagnosis with Large Language Models

  • 基于GPT-4o与检索增强生成,动态对话收集并标准化症状。
  • 准确率达90%,前3名诊断命中率100%,超越传统机器学习模型。
  • 适合需要透明、交互式医疗AI的临床场景与研究者参考。

全球医疗系统面临诊断效率低、成本高及专科医生资源不足等问题,常导致治疗延迟和不良健康结局。现有AI诊断系统多缺乏交互性与透明度,难以在以患者为中心的临床环境中有效应用。本研究提出一种基于大语言模型(LLM)的诊断聊天机器人,采用GPT-4o、检索增强生成(RAG)与可解释AI技术,通过动态对话提取并标准化患者症状,利用相似性匹配与自适应提问机制优先排序潜在诊断。结合思维链提示(Chain-of-Thought prompting),系统提供可解释的推理过程。在与朴素贝叶斯、逻辑回归、SVM、随机森林及KNN等传统模型对比中,该系统实现90%的准确率与100%的Top-3准确率,展现出在医疗AI中更透明、互动性强且临床相关的优势。

原文摘要 · Abstract (English)

Healthcare systems around the world are grappling with issues like inefficient diagnostics, rising costs, and limited access to specialists. These problems often lead to delays in treatment and poor health outcomes. Most current AI and deep learning diagnostic systems are not very interactive or transparent, making them less effective in real-world, patient-centered environments. This research introduces a diagnostic chatbot powered by a Large Language Model (LLM), using GPT-4o, Retrieval-Augmented Generation, and explainable AI techniques. The chatbot engages patients in a dynamic conversation, helping to extract and normalize symptoms while prioritizing potential diagnoses through similarity matching and adaptive questioning. With Chain-of-Thought prompting, the system also offers more transparent reasoning behind its diagnoses. When tested against traditional machine learning models like Naive Bayes, Logistic Regression, SVM, Random Forest, and KNN, the LLM-based system delivered impressive results, achieving an accuracy of 90% and Top-3 accuracy of 100%. These findings offer a promising outlook for more transparent, interactive, and clinically relevant AI in healthcare.

可解释AI医疗问答大模型应用诊断系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。