arXiv:2502.03511cs.SEcs.AI2025-02被引 2

测试ChatGPT在任务工程问题定义中的表现,发现其结果不稳定且易偏题。

An Empirical Exploration of ChatGPT's Ability to Support Problem Formulation Tasks for Mission Engineering and a Documentation of its Performance Variability

  • 用多轮测试对比ChatGPT识别利益相关方的能力
  • 对人相关方识别较好,但对外部系统和环境因素识别差
  • 输出差异大,不适合直接用于关键问题定义任务

系统工程正面临生成式AI兴起与体系化任务工程(ME)需求的双重推动。问题定义是其中核心挑战——需将模糊需求转化为可工程化的问题。本文聚焦于大型语言模型(LLM)在支持任务工程问题定义中的作用,以美国国家航空航天局(NASA)太空任务设计挑战为参考案例,评估ChatGPT-3.5在利益相关方识别任务上的表现。通过多轮并行测试与定性分析,发现模型在识别人类相关利益方方面表现尚可,但对系统外环境与外部系统因素识别能力差,且常产生过细、偏向解决方案的输出,违背问题定义初衷。更显著的是,不同运行间输出差异极大,凸显其不可靠性。研究建议:虽可辅助降低专家负担,但应以随机性视角谨慎使用,避免依赖。

原文摘要 · Abstract (English)

Systems engineering (SE) is evolving with the availability of generative artificial intelligence (AI) and the demand for a systems-of-systems perspective, formalized under the purview of mission engineering (ME) in the US Department of Defense. Formulating ME problems is challenging because they are open-ended exercises that involve translation of ill-defined problems into well-defined ones that are amenable for engineering development. It remains to be seen to which extent AI could assist problem formulation objectives. To that end, this paper explores the quality and consistency of multi-purpose Large Language Models (LLM) in supporting ME problem formulation tasks, specifically focusing on stakeholder identification. We identify a relevant reference problem, a NASA space mission design challenge, and document ChatGPT-3.5's ability to perform stakeholder identification tasks. We execute multiple parallel attempts and qualitatively evaluate LLM outputs, focusing on both their quality and variability. Our findings portray a nuanced picture. We find that the LLM performs well in identifying human-focused stakeholders but poorly in recognizing external systems and environmental factors, despite explicit efforts to account for these. Additionally, LLMs struggle with preserving the desired level of abstraction and exhibit a tendency to produce solution specific outputs that are inappropriate for problem formulation. More importantly, we document great variability among parallel threads, highlighting that LLM outputs should be used with caution, ideally by adopting a stochastic view of their abilities. Overall, our findings suggest that, while ChatGPT could reduce some expert workload, its lack of consistency and domain understanding may limit its reliability for problem formulation tasks.

任务工程LLM评测问题定义ChatGPT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。