arXiv:2603.04452cs.CLcs.AI2026-03

首个面向燃烧科学的LLM知识注入与评估框架,提升模型专业能力。

A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science

  • 构建35亿词规模多模态知识库,整合20万篇论文与40万行代码
  • 标准RAG准确率仅60%,受上下文污染限制,无法突破瓶颈
  • 提出三阶段知识注入路径,需结构化知识图谱和持续预训练

为推进燃烧科学领域基础大语言模型的发展,本研究首次提出端到端的专用模型开发框架。该框架包含一个35亿词规模、面向AI的多模态知识库,数据源自20万余篇同行评审论文、8,000篇硕博论文及约40万行燃烧计算流体动力学(CFD)代码;一个严格且高度自动化的评估基准(CombustionQA),涵盖八个子领域的436个问题;以及从轻量级检索增强生成(RAG)到知识图谱增强检索再到持续预训练的三阶段知识注入路径。我们首先定量验证第一阶段(朴素RAG)的表现,发现其准确率存在硬上限:最高仅为60%,虽远超零样本性能(23%),但仍显著低于理论上限(87%)。进一步证明该阶段性能受限于上下文污染。因此,构建领域基础模型必须依赖结构化知识图谱和持续预训练(第二、三阶段)。

原文摘要 · Abstract (English)

To advance foundation Large Language Models (LLMs) for combustion science, this study presents the first end-to-end framework for developing domain-specialized models for the combustion community. The framework comprises an AI-ready multimodal knowledge base at the 3.5 billion-token scale, extracted from over 200,000 peer-reviewed articles, 8,000 theses and dissertations, and approximately 400,000 lines of combustion CFD code; a rigorous and largely automated evaluation benchmark (CombustionQA, 436 questions across eight subfields); and a three-stage knowledge-injection pathway that progresses from lightweight retrieval-augmented generation (RAG) to knowledge-graph-enhanced retrieval and continued pretraining. We first quantitatively validate Stage 1 (naive RAG) and find a hard ceiling: standard RAG accuracy peaks at 60%, far surpassing zero-shot performance (23%) yet well below the theoretical upper bound (87%). We further demonstrate that this stage's performance is severely constrained by context contamination. Consequently, building a domain foundation model requires structured knowledge graphs and continued pretraining (Stages 2 and 3).

大模型知识注入燃烧科学评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。