专为科学推理打造的大型模型,能高效筛选电池电解液候选分子。
OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery
- 基于科学文献预训练+任务指令微调+逻辑推理知识蒸馏
- 在电池材料筛选任务中表现超越同类模型,接近顶尖水平
- 适合科研人员和工业界做科学发现与分子设计
大型语言模型在推动科学知识和应对复杂挑战方面展现出巨大潜力。本文提出OmniScience,一种面向通用科学领域的专用大推理模型,通过三个关键组件构建:(1) 在精心筛选的科学文献语料上进行领域自适应预训练;(2) 在专用数据集上进行指令微调以引导模型完成领域特定任务;(3) 通过基于推理的知识蒸馏微调显著提升其生成上下文相关且逻辑严谨回答的能力。我们通过开发一个分子筛选代理,高效评估候选电解液溶剂或添加剂。综合评估显示,OmniScience在GPQA Diamond和领域专用电池基准上表现与当前最优大推理模型相当,同时优于所有同参数量级的公开推理与非推理模型。消融实验进一步证明,领域自适应预训练和基于推理的知识蒸馏对性能提升至关重要。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable potential in advancing scientific knowledge and addressing complex challenges. In this work, we introduce OmniScience, a specialized large reasoning model for general science, developed through three key components: (1) domain adaptive pretraining on a carefully curated corpus of scientific literature, (2) instruction tuning on a specialized dataset to guide the model in following domain-specific tasks, and (3) reasoning-based knowledge distillation through fine-tuning to significantly enhance its ability to generate contextually relevant and logically sound responses. We demonstrate the versatility of OmniScience by developing a battery agent that efficiently ranks molecules as potential electrolyte solvents or additives. Comprehensive evaluations reveal that OmniScience is competitive with state-of-the-art large reasoning models on the GPQA Diamond and domain-specific battery benchmarks, while outperforming all public reasoning and non-reasoning models with similar parameter counts. We further demonstrate via ablation experiments that domain adaptive pretraining and reasoning-based knowledge distillation are critical to attain our performance levels, across benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。