专为医学领域打造的大模型,训练更高效,表现更专业。
Baichuan-M1: Pushing the Medical Capability of Large Language Models
- 从头训练,专注提升医学能力,非简单微调通用模型
- 使用20万亿词训练,医学任务表现超越通用模型
- 开源140亿参数版本,适合医疗研究与应用开发
当前大型语言模型多面向通用场景,而医疗等垂直领域的专用模型仍较稀缺。由于医学知识复杂且高质量数据有限,构建高效实用的医学大模型面临挑战。为此,我们提出Baichuan-M1,一系列专为医疗应用优化的大语言模型。不同于仅在现有模型上继续预训练或后训练的方法,Baichuan-M1从零开始训练,重点强化医学能力。模型基于20万亿个令牌训练,采用多种有效训练策略,在保持通用能力的同时显著提升医学专业性。实验表明,Baichuan-M1不仅在数学、编程等通用任务中表现优异,更在多个医学领域实现领先性能。我们已开源Baichuan-M1-14B,可通过指定链接获取。
原文摘要 · Abstract (English)
The current generation of large language models (LLMs) is typically designed for broad, general-purpose applications, while domain-specific LLMs, especially in vertical fields like medicine, remain relatively scarce. In particular, the development of highly efficient and practical LLMs for the medical domain is challenging due to the complexity of medical knowledge and the limited availability of high-quality data. To bridge this gap, we introduce Baichuan-M1, a series of large language models specifically optimized for medical applications. Unlike traditional approaches that simply continue pretraining on existing models or apply post-training to a general base model, Baichuan-M1 is trained from scratch with a dedicated focus on enhancing medical capabilities. Our model is trained on 20 trillion tokens and incorporates a range of effective training methods that strike a balance between general capabilities and medical expertise. As a result, Baichuan-M1 not only performs strongly across general domains such as mathematics and coding but also excels in specialized medical fields. We have open-sourced Baichuan-M1-14B, a mini version of our model, which can be accessed through the following links.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。