arXiv:2504.18858cs.SEcs.AI2025-04被引 3

ChatGPT在不同领域和开发阶段错误率差异大,需警惕盲目信任。

Why you shouldn't fully trust ChatGPT: A synthesis of this AI tool's error rates across disciplines and the software engineering lifecycle

  • 综合多源数据,按领域和软件开发阶段分析错误率
  • 医疗领域错误率高达83%,编程调试仍超50%出错
  • GPT-4比GPT-3.5更可靠,但关键任务仍需人工审核

ChatGPT等大语言模型在医疗、商业、工程及软件工程(SE)等领域广泛应用。本研究通过多源文献综述(MLR),整合截至2025年的学术研究、报告与基准测试数据,量化分析其在各领域及软件开发生命周期(SDLC)阶段的错误率。涵盖事实性、推理、编码与解释性错误,按领域与开发阶段分组,并用箱线图展示分布。结果显示:医疗领域错误率8%-83%;商业与经济领域从GPT-3.5的约50%降至GPT-4的15%-20%;工程任务平均20%-30%;编程成功率87.5%,但复杂调试错误率仍超50%。在软件工程中,需求与设计阶段错误率较低(~5%-20%),而编码、测试与维护阶段波动大(10%-50%)。从GPT-3.5升级至GPT-4显著提升可靠性。结论指出:尽管性能改善,但错误率在不同场景下仍不可忽视,完全依赖存在风险,尤其在关键领域,需持续评估与人工验证以保障可信度。

原文摘要 · Abstract (English)

Context: ChatGPT and other large language models (LLMs) are widely used across healthcare, business, economics, engineering, and software engineering (SE). Despite their popularity, concerns persist about their reliability, especially their error rates across domains and the software development lifecycle (SDLC). Objective: This study synthesizes and quantifies ChatGPT's reported error rates across major domains and SE tasks aligned with SDLC phases. It provides an evidence-based view of where ChatGPT excels, where it fails, and how reliability varies by task, domain, and model version (GPT-3.5, GPT-4, GPT-4-turbo, GPT-4o). Method: A Multivocal Literature Review (MLR) was conducted, gathering data from academic studies, reports, benchmarks, and grey literature up to 2025. Factual, reasoning, coding, and interpretive errors were considered. Data were grouped by domain and SE phase and visualized using boxplots to show error distributions. Results: Error rates vary across domains and versions. In healthcare, rates ranged from 8% to 83%. Business and economics saw error rates drop from ~50% with GPT-3.5 to 15-20% with GPT-4. Engineering tasks averaged 20-30%. Programming success reached 87.5%, though complex debugging still showed over 50% errors. In SE, requirements and design phases showed lower error rates (~5-20%), while coding, testing, and maintenance phases had higher variability (10-50%). Upgrades from GPT-3.5 to GPT-4 improved reliability. Conclusion: Despite improvements, ChatGPT still exhibits non-negligible error rates varying by domain, task, and SDLC phase. Full reliance without human oversight remains risky, especially in critical settings. Continuous evaluation and critical validation are essential to ensure reliability and trustworthiness.

大模型评估错误率分析软件工程可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。