arXiv:2412.16336cs.SEcs.AI2024-12被引 3

实证发现多数深度学习故障基准数据不真实,仅18.5%符合真实条件。

Real Faults in Deep Learning Fault Benchmarks: How Real Are They?

  • 人工分析490个故障,筛选出314个可研究样本
  • 仅18.5%故障满足真实性标准,52%可复现
  • 适合关注模型测试可信度的研究者

随着深度学习系统应用日益广泛,越来越多方法被提出用于测试、定位和修复系统缺陷。评估这些方法有效性的最佳依据是其能否检测、定位并修复真实故障。为此,研究社区构建了多个深度学习系统的真实故障基准。本文对五个基准中的490个故障进行人工分析,识别出314个符合条件的故障。研究重点考察故障与其原始来源的对应程度、故障类型分布以及可复现性。结果表明,仅有18.5%的故障满足我们定义的真实性条件;在尝试复现时,成功案例占52%。

原文摘要 · Abstract (English)

As the adoption of Deep Learning (DL) systems continues to rise, an increasing number of approaches are being proposed to test these systems, localise faults within them, and repair those faults. The best attestation of effectiveness for such techniques is an evaluation that showcases their capability to detect, localise and fix real faults. To facilitate these evaluations, the research community has collected multiple benchmarks of real faults in DL systems. In this work, we perform a manual analysis of 490 faults from five different benchmarks and identify that 314 of them are eligible for our study. Our investigation focuses specifically on how well the bugs correspond to the sources they were extracted from, which fault types are represented, and whether the bugs are reproducible. Our findings indicate that only 18.5% of the faults satisfy our realism conditions. Our attempts to reproduce these faults were successful only in 52% of cases.

深度学习故障测试基准验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。