构建中文医疗对话数据集,提升智能问诊准确率
Building a Chinese Medical Dialogue System: Integrating Large-scale Corpora and Novel Models
- 自建大规模中文医疗对话数据集,解决数据稀缺问题
- 融合BERT与GPT的模型架构,问诊准确率显著提升
- 适合医疗AI研究者及智能问诊系统开发者参考
全球新冠疫情凸显传统医疗体系短板,加速在线医疗服务发展,尤其在医疗分诊与咨询领域。然而现有研究面临两大挑战:一是因隐私限制,缺乏大规模公开的领域专用医疗数据集,现有数据规模小且病种有限,制约基于预训练语言模型(PLMs)的分诊效果;二是现有方法缺乏医学知识,难以准确理解患者-医生对话中的专业术语和表达。为此,我们构建了大规模中文医疗对话语料库(LCMDC),缓解数据不足问题。同时提出一种结合BERT监督学习与提示学习的新型分诊系统,以及基于GPT的医疗咨询模型。为增强领域知识获取能力,我们使用自建背景语料对PLMs进行预训练。在LCMDC上的实验结果验证了所提系统的有效性。
原文摘要 · Abstract (English)
The global COVID-19 pandemic underscored major deficiencies in traditional healthcare systems, hastening the advancement of online medical services, especially in medical triage and consultation. However, existing studies face two main challenges. First, the scarcity of large-scale, publicly available, domain-specific medical datasets due to privacy concerns, with current datasets being small and limited to a few diseases, limiting the effectiveness of triage methods based on Pre-trained Language Models (PLMs). Second, existing methods lack medical knowledge and struggle to accurately understand professional terms and expressions in patient-doctor consultations. To overcome these obstacles, we construct the Large-scale Chinese Medical Dialogue Corpora (LCMDC), thereby addressing the data shortage in this field. Moreover, we further propose a novel triage system that combines BERT-based supervised learning with prompt learning, as well as a GPT-based medical consultation model. To enhance domain knowledge acquisition, we pre-trained PLMs using our self-constructed background corpus. Experimental results on the LCMDC demonstrate the efficacy of our proposed systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。