arXiv:2603.19268cs.CLcs.AI2026-03

为燃烧科学定制大模型,解决幻觉问题并提升物理规律遵循能力。

Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization

  • 构建全流程增强框架,融合语料生成、增量预训练与可验证强化学习。
  • 在燃烧推理任务上超越闭源主流模型和检索增强方法。
  • 提供专用评测基准FlameBench,适合科研智能体开发人员使用。

面向专业领域任务适配与能力提升的大语言模型展现出巨大应用潜力。然而,对于燃烧科学等复杂物理系统,通用大模型常因领域知识不足及无法遵守物理守恒定律而产生严重幻觉。为此,本文提出首个专用于燃烧科学的全栈领域增强大模型流程,涵盖自动化领域语料构建、增量预训练、指令微调与可验证奖励驱动的强化学习。该流程确保模型真正内化物理规律,而非仅学习文本统计模式。同时发布专门针对燃烧科学复杂推理任务的标准化评测基准FlameBench。实验表明,本工作所开发模型在燃烧科学推理任务上显著优于当前最先进的闭源通用模型与传统检索增强生成方法。本研究为后续具备可靠科学推理能力的领域专用科研智能体发展奠定了坚实的技术与资源基础。

原文摘要 · Abstract (English)

Large language models (LLMs) in the direction of task adaptation and capability enhancement for professional fields demonstrate significant application potential. Nevertheless, for complex physical systems such as combustion science, general-purpose LLMs often generate severe hallucinations due to insufficient domain knowledge and the inability to adhere to physical conservation laws. To address this issue, we propose the first full-stack domain-enhanced LLM workflow tailored for the field of combustion science, which integrates automated domain corpus construction, incremental pre-training, instruction fine-tuning, and verifiable reward-based reinforcement learning. This workflow ensures that the model truly internalizes physical laws rather than merely learning textual statistical patterns. We also release FlameBench, a standardized evaluation benchmark specifically designed for complex reasoning tasks in combustion science. Experimental results demonstrate that the model developed in this work significantly outperforms state-of-the-art general-purpose closed-source models and traditional retrieval-augmented generation methods on combustion science reasoning tasks. This work lays a solid technical and resource foundation for the subsequent development of domain-specific scientific research agents with reliable scientific reasoning capabilities.

燃烧科学大模型领域增强科学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。