用分阶段推理提升医疗对话模型的诊断能力
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
- 分三阶段模拟医生问诊:症状采集、病史获取、因果图构建
- 诊断准确率高出基线5.18%,症状回忆提升超30%
- 适合需要透明可解释医疗辅助的临床场景
尽管大型语言模型(LLMs)表现出色,现有对话式健康助手(CHAs)仍僵化脆弱,缺乏自适应多轮推理、症状澄清与透明决策能力,限制了其在临床诊断中的实际应用。我们提出DocCHA,一种基于置信度感知的模块化框架,通过将诊断过程分解为三个阶段:(1) 症状采集,(2) 病史获取,(3) 因果图构建。各模块使用可解释的置信度分数指导自适应提问、优先澄清关键信息、优化薄弱推理链。在两个真实世界中文咨询数据集(IMCS21, DX)上评估,DocCHA持续优于强提示基线(GPT-3.5, GPT-4o, LLaMA-3),诊断准确率最高提升5.18%,症状回忆率提高超30%,仅小幅增加对话轮次。结果表明,DocCHA能实现结构化、透明且高效的诊断对话,为多语言和资源受限环境中的可信大模型临床助手铺平道路。
原文摘要 · Abstract (English)
Despite the impressive capabilities of Large Language Models (LLMs), existing Conversational Health Agents (CHAs) remain static and brittle, incapable of adaptive multi-turn reasoning, symptom clarification, or transparent decision-making. This hinders their real-world applicability in clinical diagnosis, where iterative and structured dialogue is essential. We propose DocCHA, a confidence-aware, modular framework that emulates clinical reasoning by decomposing the diagnostic process into three stages: (1) symptom elicitation, (2) history acquisition, and (3) causal graph construction. Each module uses interpretable confidence scores to guide adaptive questioning, prioritize informative clarifications, and refine weak reasoning links. Evaluated on two real-world Chinese consultation datasets (IMCS21, DX), DocCHA consistently outperforms strong prompting-based LLM baselines (GPT-3.5, GPT-4o, LLaMA-3), achieving up to 5.18 percent higher diagnostic accuracy and over 30 percent improvement in symptom recall, with only modest increase in dialogue turns. These results demonstrate the effectiveness of DocCHA in enabling structured, transparent, and efficient diagnostic conversations -- paving the way for trustworthy LLM-powered clinical assistants in multilingual and resource-constrained settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。