用多源知识重建和评估临床问诊对话,提升大模型诊疗能力。
CliniChat: A Multi-Source Knowledge-Driven Framework for Clinical Interview Dialogue Reconstruction and Evaluation
- 融合三类知识源,将病历记录转化为专业且共情的问诊对话。
- 构建自动评估体系,使模型可像专家一样评价问诊表现。
- 推出高质量合成数据集与专用模型,显著提升问诊能力。
大语言模型在辅助临床问诊方面前景广阔,但缺乏高质量对话数据和统一评估方法严重制约其发展。为此,我们提出 Clinichat 框架,通过两个模块实现问诊对话的重建与评估:Clini-Recon 利用三类知识源,将临床笔记转化为系统、专业且富有同理心的对话;Clini-Eval 构建综合评估指标体系与两阶段自动化评估方法,使模型能像专家一样评估问诊表现。我们贡献了 MedQA-Dialog 高质量合成对话数据集及 CliniChatGLM 专用模型。实验表明,CliniChatGLM 在问诊能力上实现全面升级,尤其在病史采集任务中达到领先水平。
原文摘要 · Abstract (English)
Large language models (LLMs) hold great promise for assisting clinical interviews due to their fluent interactive capabilities and extensive medical knowledge. However, the lack of high-quality interview dialogue data and widely accepted evaluation methods has significantly impeded this process. So we propose CliniChat, a framework that integrates multi-source knowledge to enable LLMs to simulate real-world clinical interviews. It consists of two modules: Clini-Recon and Clini-Eval, each responsible for reconstructing and evaluating interview dialogues, respectively. By incorporating three sources of knowledge, Clini-Recon transforms clinical notes into systematic, professional, and empathetic interview dialogues. Clini-Eval combines a comprehensive evaluation metric system with a two-phase automatic evaluation approach, enabling LLMs to assess interview performance like experts. We contribute MedQA-Dialog, a high-quality synthetic interview dialogue dataset, and CliniChatGLM, a model specialized for clinical interviews. Experimental results demonstrate that CliniChatGLM's interview capabilities undergo a comprehensive upgrade, particularly in history-taking, achieving state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。