CareBot是首个全流程开源的双语医学大模型,提升医疗问答与对话能力。
CareBot: A Pioneering Full-Process Open-Source Medical Language Model
- 采用两阶段持续预训练,逐步强化医学知识
- 构建高质量双语数据集,对话质量显著提升
- 适合医疗AI研究者与开发者使用
近年来,闭源与开源大模型在通用领域表现超越人类,但在医学等专业领域,尤其在开源社区中仍表现不佳,源于医学知识的复杂性。本文提出CareBot,一个双语医学大模型,融合连续预训练(CPT)、监督微调(SFT)和基于人类反馈的强化学习(RLHF)。创新性地设计两阶段CPT方法:稳定预训练与增强预训练,有效弥合通用与专业数据差距,实现知识渐进式增强。引入DataRater模型评估预训练数据质量,确保数据准确相关。SFT阶段构建大规模多样双语数据集,并开发ConFilter指标,显著提升多轮对话质量。在中英文医疗基准上的严格评估表明,CareBot在医疗咨询与教育任务中表现优异。这些进展不仅突破当前医学大模型局限,更确立了开源医学模型开发新标准。后续将开源数据集与模型,为研究社区提供宝贵资源。
原文摘要 · Abstract (English)
Recently, both closed-source LLMs and open-source communities have made significant strides, outperforming humans in various general domains. However, their performance in specific professional domains such as medicine, especially within the open-source community, remains suboptimal due to the complexity of medical knowledge. In this paper, we propose CareBot, a bilingual medical LLM, which leverages a comprehensive approach integrating continuous pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning with human feedback (RLHF). Our novel two-stage CPT method, comprising Stable CPT and Boost CPT, effectively bridges the gap between general and domain-specific data, facilitating a smooth transition from pre-training to fine-tuning and enhancing domain knowledge progressively. We also introduce DataRater, a model designed to assess data quality during CPT, ensuring that the training data is both accurate and relevant. For SFT, we develope a large and diverse bilingual dataset, along with ConFilter, a metric to enhance multi-turn dialogue quality, which is crucial to improving the model's ability to handle more complex dialogues. The combination of high-quality data sources and innovative techniques significantly improves CareBot's performance across a range of medical applications. Our rigorous evaluations on Chinese and English benchmarks confirm CareBot's effectiveness in medical consultation and education. These advancements not only address current limitations in medical LLMs but also set a new standard for developing effective and reliable open-source models in the medical domain. We will open-source the datasets and models later, contributing valuable resources to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。