arXiv:2502.09690cs.CLcs.AI2025-02被引 12

大模型生成的系统工程文档看似专业,实则暗藏致命缺陷。

Trust at Your Own Peril: A Mixed Methods Exploration of the Ability of Large Language Models to Generate Expert-Like Systems Engineering Artifacts and a Characterization of Failure Modes

  • 用提示工程让大模型生成系统工程文档,不微调直接测试基础能力
  • 语言模型生成内容与专家文档在统计上几乎无法区分,但质量存隐患
  • 发现三大隐蔽错误:过早定义需求、无依据数值、过度详细设计

多用途大语言模型(LLMs)作为生成式人工智能的代表,近年来进展显著。尽管人们期望其能协助系统工程(SE)任务,但系统工程的跨学科性与复杂性,以及对深层领域知识和操作背景的整合需求,使大模型生成高质量工程文档的能力备受质疑——尤其是其训练数据来自互联网公开信息。为此,我们开展实证研究:以人工专家生成的系统工程文档为基准,经解析后通过提示工程输入不同大模型,生成典型系统工程文档片段。该过程未进行微调或校准,旨在评估大模型的基线性能。采用两步混合方法对比生成结果与基准:首先,使用自然语言处理算法定量比较,发现精心提示下,当前顶尖模型生成的内容与人类专家基准难以区分;其次,通过定性深度分析,揭示两者虽表面相似,但生成内容存在严重失效模式,包括:过早定义需求、缺乏依据的数值估算、过度详细化倾向。研究警示:系统工程领域应警惕依赖大模型生成的反馈,尤其是在未加验证的情况下。

原文摘要 · Abstract (English)

Multi-purpose Large Language Models (LLMs), a subset of generative Artificial Intelligence (AI), have recently made significant progress. While expectations for LLMs to assist systems engineering (SE) tasks are paramount; the interdisciplinary and complex nature of systems, along with the need to synthesize deep-domain knowledge and operational context, raise questions regarding the efficacy of LLMs to generate SE artifacts, particularly given that they are trained using data that is broadly available on the internet. To that end, we present results from an empirical exploration, where a human expert-generated SE artifact was taken as a benchmark, parsed, and fed into various LLMs through prompt engineering to generate segments of typical SE artifacts. This procedure was applied without any fine-tuning or calibration to document baseline LLM performance. We then adopted a two-fold mixed-methods approach to compare AI generated artifacts against the benchmark. First, we quantitatively compare the artifacts using natural language processing algorithms and find that when prompted carefully, the state-of-the-art algorithms cannot differentiate AI-generated artifacts from the human-expert benchmark. Second, we conduct a qualitative deep dive to investigate how they differ in terms of quality. We document that while the two-material appear very similar, AI generated artifacts exhibit serious failure modes that could be difficult to detect. We characterize these as: premature requirements definition, unsubstantiated numerical estimates, and propensity to overspecify. We contend that this study tells a cautionary tale about why the SE community must be more cautious adopting AI suggested feedback, at least when generated by multi-purpose LLMs.

系统工程大模型风险生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。