SciDFM用专家混合架构,让大模型懂化学分子和蛋白质序列
SciDFM: A Large Language Model with Mixture-of-Experts for Science
- 采用专家混合架构,从头训练以融合跨学科科学知识
- 在通用科学评测上表现优异,特定领域任务达同类模型最佳
- 适合需要理解分子与生物序列的科研人员使用
近期,大语言模型在辅助科学发现方面受到广泛关注。然而,多数模型仅关注通用科学,缺乏对化学分子、氨基酸序列等专业领域的知识。为此,我们提出SciDFM,一种基于专家混合(MoE)的大语言模型,从零训练,具备大学水平的科学推理能力,并能理解分子结构与氨基酸序列。我们构建了涵盖多学科科学论文与书籍的大规模语料库,还整合了领域专用数据库数据。通过大量指令微调,显著提升其下游任务表现。实验表明,SciDFM在SciEval和SciQ等通用科学评测中表现强劲,在同规模模型中于特定领域任务达到最先进水平。我们进一步分析专家层,发现不同学科数据下专家选择策略存在差异。为促进社区共享,我们已在Hugging Face开源SciDFM-MoE-A5.6B-v1.0。
原文摘要 · Abstract (English)
Recently, there has been a significant upsurge of interest in leveraging large language models (LLMs) to assist scientific discovery. However, most LLMs only focus on general science, while they lack domain-specific knowledge, such as chemical molecules and amino acid sequences. To bridge these gaps, we introduce SciDFM, a mixture-of-experts LLM, which is trained from scratch and is able to conduct college-level scientific reasoning and understand molecules and amino acid sequences. We collect a large-scale training corpus containing numerous scientific papers and books from different disciplines as well as data from domain-specific databases. We further fine-tune the pre-trained model on lots of instruction data to improve performances on downstream benchmarks. From experiment results, we show that SciDFM achieves strong performance on general scientific benchmarks such as SciEval and SciQ, and it reaches a SOTA performance on domain-specific benchmarks among models of similar size. We further analyze the expert layers and show that the results of expert selection vary with data from different disciplines. To benefit the broader research community, we open-source SciDFM at https://huggingface.co/OpenDFM/SciDFM-MoE-A5.6B-v1.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。