arXiv:2501.17183cs.CLcs.AI2025-01被引 10

为航天制造领域量身打造LLM评估体系,解决模型幻觉风险

LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering

  • 从专业教材提取关键知识,生成多正确答案的多选题
  • 测试多个LLM在航天制造题上的准确率,发现表现亟待提升
  • 适合关注工业AI安全、高精度模型评估的研究者

航天制造对技术参数精度要求极高。尽管大语言模型(如GPT-4和QWen)在自然语言处理中表现优异,引发其在工艺设计、材料选择和工具信息检索等任务中的应用兴趣,但其在专业领域易产生“幻觉”,输出错误信息,威胁航天产品质量与飞行安全。本文提出一套面向航天制造领域的LLM评估指标,通过深度分析经典教材与规范,提取关键知识,利用LLM生成技术构建多难度、多正确答案的多项选择题。随后采用不同模型作答并记录准确率。实验表明,当前LLM在航天专业知识上的能力亟需提升。该研究为大模型在航天制造中的安全应用提供了理论基础与实践指导,填补了该领域的关键空白。

原文摘要 · Abstract (English)

Aerospace manufacturing demands exceptionally high precision in technical parameters. The remarkable performance of Large Language Models (LLMs), such as GPT-4 and QWen, in Natural Language Processing has sparked industry interest in their application to tasks including process design, material selection, and tool information retrieval. However, LLMs are prone to generating "hallucinations" in specialized domains, producing inaccurate or false information that poses significant risks to the quality of aerospace products and flight safety. This paper introduces a set of evaluation metrics tailored for LLMs in aerospace manufacturing, aiming to assess their accuracy by analyzing their performance in answering questions grounded in professional knowledge. Firstly, key information is extracted through in-depth textual analysis of classic aerospace manufacturing textbooks and guidelines. Subsequently, utilizing LLM generation techniques, we meticulously construct multiple-choice questions with multiple correct answers of varying difficulty. Following this, different LLM models are employed to answer these questions, and their accuracy is recorded. Experimental results demonstrate that the capabilities of LLMs in aerospace professional knowledge are in urgent need of improvement. This study provides a theoretical foundation and practical guidance for the application of LLMs in aerospace manufacturing, addressing a critical gap in the field.

LLM评估航天制造幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。