arXiv:2509.02227cs.IRcs.AI2025-09

用大模型自动提取加速器技术文档中的专业知识,防止经验流失。

Application Of Large Language Models For The Extraction Of Information From Particle Accelerator Technical Documentation

  • 用大语言模型解析粒子加速器的老旧技术文档,提取关键信息。
  • 模型能有效摘要和组织知识,降低资深人员退休带来的风险。
  • 适合科研机构和高精尖领域知识管理,尤其对经验传承有帮助。

大量老旧粒子加速器的技术文档,加上资深人员陆续退休,凸显出高效保存与传递专业经验的迫切需求。本文探索使用大语言模型(LLMs)自动化并增强从粒子加速器技术文档中提取信息的能力。通过利用LLMs,旨在解决知识留存难题,实现对嵌入在历史文档中的领域专长的检索。我们展示了将LLMs适配到该专业领域的初步成果。评估表明,LLMs在信息提取、摘要生成和知识组织方面均表现出色,显著降低了因人员更替而丢失宝贵见解的风险。此外,我们讨论了当前LLMs在可解释性及处理罕见领域术语方面的局限性,并提出了改进策略。本研究凸显了LLMs在保存机构知识和保障高度专业化领域持续性方面的重要潜力。

原文摘要 · Abstract (English)

The large set of technical documentation of legacy accelerator systems, coupled with the retirement of experienced personnel, underscores the urgent need for efficient methods to preserve and transfer specialized knowledge. This paper explores the application of large language models (LLMs), to automate and enhance the extraction of information from particle accelerator technical documents. By exploiting LLMs, we aim to address the challenges of knowledge retention, enabling the retrieval of domain expertise embedded in legacy documentation. We present initial results of adapting LLMs to this specialized domain. Our evaluation demonstrates the effectiveness of LLMs in extracting, summarizing, and organizing knowledge, significantly reducing the risk of losing valuable insights as personnel retire. Furthermore, we discuss the limitations of current LLMs, such as interpretability and handling of rare domain-specific terms, and propose strategies for improvement. This work highlights the potential of LLMs to play a pivotal role in preserving institutional knowledge and ensuring continuity in highly specialized fields.

大模型知识提取加速器文献理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。