SciGPT专为科学文献理解设计,提升跨学科研究效率。
SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery
- 基于Qwen3架构,用轻量领域蒸馏优化性能与效率
- 长文档推理内存降低55%,支持32000词上下文分析
- 融合领域知识图谱,提升跨学科任务泛化能力
科学文献呈指数增长,研究人员面临知识整合瓶颈。通用大模型虽具文本处理潜力,却难以捕捉科学领域的专业特征(如技术术语、方法严谨性),在复杂科学任务中表现有限。为此,本文提出面向科学文献理解的领域适配基础模型SciGPT,以及专用于评估科学大模型的开源基准ScienceBench。SciGPT基于Qwen3架构,包含三项创新:(1) 采用两阶段低代价领域蒸馏策略,兼顾性能与效率;(2) 引入稀疏专家混合注意力机制,在32,000词长文档推理中将内存消耗降低55%;(3) 融合领域本体的知识增强适配,弥合跨学科知识鸿沟。在ScienceBench上的实验表明,SciGPT在序列标注、生成和推理等核心科学任务中优于GPT-4o,且在未见任务中表现出强鲁棒性,验证其在人工智能辅助科学发现中的潜力。
原文摘要 · Abstract (English)
Scientific literature is growing exponentially, creating a critical bottleneck for researchers to efficiently synthesize knowledge. While general-purpose Large Language Models (LLMs) show potential in text processing, they often fail to capture scientific domain-specific nuances (e.g., technical jargon, methodological rigor) and struggle with complex scientific tasks, limiting their utility for interdisciplinary research. To address these gaps, this paper presents SciGPT, a domain-adapted foundation model for scientific literature understanding and ScienceBench, an open source benchmark tailored to evaluate scientific LLMs. Built on the Qwen3 architecture, SciGPT incorporates three key innovations: (1) low-cost domain distillation via a two-stage pipeline to balance performance and efficiency; (2) a Sparse Mixture-of-Experts (SMoE) attention mechanism that cuts memory consumption by 55\% for 32,000-token long-document reasoning; and (3) knowledge-aware adaptation integrating domain ontologies to bridge interdisciplinary knowledge gaps. Experimental results on ScienceBench show that SciGPT outperforms GPT-4o in core scientific tasks including sequence labeling, generation, and inference. It also exhibits strong robustness in unseen scientific tasks, validating its potential to facilitate AI-augmented scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。