首份系统分析代码审查自动化研究,揭示标准缺失与方法挑战
Previously on... Automating Code Review
- 梳理691篇文献,聚焦24项核心任务的定义与评估差异
- 发现48种指标组合,其中22种仅在原论文中使用
- 提出时间偏差等关键问题,推动评估标准化
现代代码审查(MCR)是软件工程中的标准实践,但耗费大量时间和资源。近年来,越来越多研究尝试用机器学习(ML)和深度学习(DL)自动化核心审查任务,但任务定义、数据集和评估方式差异显著。本研究首次对MCR自动化研究进行全面分析,旨在刻画领域发展脉络、规范学习任务、揭示方法论挑战,并提出可操作建议。我们系统调研了691篇文献,识别出2015年5月至2024年4月间发表的24项相关研究。每项研究从任务、模型、指标、基线、结果、有效性问题及资源可用性等方面进行分析。结果显示,存在显著的标准化潜力:共发现48种任务-指标组合,其中22种为论文独有;数据集复用率较低。研究还揭示了如时间偏差威胁等未被充分关注的问题。本工作为领域提供清晰全景图,支持新研究设计,避免常见陷阱,促进评估实践规范化。
原文摘要 · Abstract (English)
Modern Code Review (MCR) is a standard practice in software engineering, yet it demands substantial time and resource investments. Recent research has increasingly explored automating core review tasks using machine learning (ML) and deep learning (DL). As a result, there is substantial variability in task definitions, datasets, and evaluation procedures. This study provides the first comprehensive analysis of MCR automation research, aiming to characterize the field's evolution, formalize learning tasks, highlight methodological challenges, and offer actionable recommendations to guide future research. Focusing on the primary code review tasks, we systematically surveyed 691 publications and identified 24 relevant studies published between May 2015 and April 2024. Each study was analyzed in terms of tasks, models, metrics, baselines, results, validity concerns, and artifact availability. In particular, our analysis reveals significant potential for standardization, including 48 task metric combinations, 22 of which were unique to their original paper, and limited dataset reuse. We highlight challenges and derive concrete recommendations for examples such as the temporal bias threat, which are rarely addressed so far. Our work contributes to a clearer overview of the field, supports the framing of new research, helps to avoid pitfalls, and promotes greater standardization in evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。