arXiv:2510.04997cs.SEcs.AI2025-10被引 1

用大模型自动分析软件缺陷,效率提升数倍。

AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis

  • 拆解缺陷研究为三阶段,用大模型自动化处理
  • 处理3829个缺陷仅需约两小时,人工需数周
  • 适合想快速开展软件缺陷实证研究的团队

理解软件缺陷对软件开发与维护的实证研究至关重要。传统缺陷分析依赖多步人工操作,如缺陷收集、筛选和手动调查,耗时费力,难以支撑复杂关键系统的大规模研究,制约了实证研究的迭代速度。本文将实证软件缺陷研究分解为三个关键阶段:(1)研究目标定义,(2)数据准备,(3)缺陷分析,并首次探索使用大语言模型(LLMs)对开源软件缺陷进行分析。我们在来自高质量实证研究的3,829个软件缺陷上进行了评估。结果表明,大模型可显著提升分析效率,平均处理时间约为两小时,远低于以往数周的人工工作量。最后,我们提出一份详细研究计划,指明大模型在推动实证缺陷研究中的潜力及实现端到端自动化所面临的开放挑战。

原文摘要 · Abstract (English)

Understanding software faults is essential for empirical research in software development and maintenance. However, traditional fault analysis, while valuable, typically involves multiple expert-driven steps such as collecting potential faults, filtering, and manual investigation. These processes are both labor-intensive and time-consuming, creating bottlenecks that hinder large-scale fault studies in complex yet critical software systems and slow the pace of iterative empirical research. In this paper, we decompose the process of empirical software fault study into three key phases: (1) research objective definition, (2) data preparation, and (3) fault analysis, and we conduct an initial exploration study of applying Large Language Models (LLMs) for fault analysis of open-source software. Specifically, we perform the evaluation on 3,829 software faults drawn from a high-quality empirical study. Our results show that LLMs can substantially improve efficiency in fault analysis, with an average processing time of about two hours, compared to the weeks of manual effort typically required. We conclude by outlining a detailed research plan that highlights both the potential of LLMs for advancing empirical fault studies and the open challenges that required be addressed to achieve fully automated, end-to-end software fault analysis.

大模型软件缺陷自动化研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。