arXiv:2506.08750cs.CL2025-06被引 1

用合成数据让核工业文本数据可用,提升大模型应用能力

Unlocking the Potential of Large Language Models in the Nuclear Industry with Synthetic Data

  • 用大模型自动将核工业文本转为问答对
  • 解决数据少且隐私敏感的难题
  • 适合核能领域知识工程与智能检索研究者

核工业拥有大量未结构化的文本信息,但难以直接用于大语言模型(LLM)训练、微调和评估,因这些任务需要干净的结构化问答对。本文探讨如何通过合成数据生成技术填补这一空白,实现核领域鲁棒大模型的开发。针对核工业存在的数据稀缺和隐私问题,合成数据方法可将现有文本转化为可用的问答对。该方法利用大模型分析文本、提取关键信息、生成相关问题,并评估合成数据集质量。通过释放大模型在核领域的潜力,合成数据有助于提升信息检索效率、促进知识共享,并支持更科学的决策制定。

原文摘要 · Abstract (English)

The nuclear industry possesses a wealth of valuable information locked away in unstructured text data. This data, however, is not readily usable for advanced Large Language Model (LLM) applications that require clean, structured question-answer pairs for tasks like model training, fine-tuning, and evaluation. This paper explores how synthetic data generation can bridge this gap, enabling the development of robust LLMs for the nuclear domain. We discuss the challenges of data scarcity and privacy concerns inherent in the nuclear industry and how synthetic data provides a solution by transforming existing text data into usable Q&A pairs. This approach leverages LLMs to analyze text, extract key information, generate relevant questions, and evaluate the quality of the resulting synthetic dataset. By unlocking the potential of LLMs in the nuclear industry, synthetic data can pave the way for improved information retrieval, enhanced knowledge sharing, and more informed decision-making in this critical sector.

大模型合成数据核工业知识抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。