arXiv:2512.23745cs.LGcs.SE2025-12被引 2

对比16种深度学习模型修复方法,发现模型级修复最有效但难兼顾其他性能。

A Comprehensive Study of Deep Learning Model Fixing Approaches

  • 系统评估16种模型修复方法,涵盖模型、层和神经元三个层级。
  • 模型级修复在修正错误上效果最好,但难以同时提升准确率并保持鲁棒性。
  • 研究揭示修复副作用普遍,适合关注模型可靠性与安全性的开发者参考。

深度学习已在自动驾驶、智能医疗、辅助编程等工业领域广泛应用,但其故障可能带来重大风险。为此,学界提出了众多模型修复方法。本文对16种前沿的深度学习模型修复技术(包括模型级、层级、神经元级)进行了大规模实证研究,全面评估其修复效果及对鲁棒性、公平性、后向兼容性等关键属性的影响。实验在统一框架下使用多样化的数据集、模型架构和应用领域进行,确保评估的全面性与公平性。研究发现:模型级修复在修复有效性上表现最优;但无一种方法能同时实现最佳修复效果、提升准确率并维持所有其他属性。该结果表明,未来研究应更关注修复带来的副作用。这些发现为产业界与学术界提供了重要启示。

原文摘要 · Abstract (English)

Deep Learning (DL) has been widely adopted in diverse industrial domains, including autonomous driving, intelligent healthcare, and aided programming. Like traditional software, DL systems are also prone to faults, whose malfunctioning may expose users to significant risks. Consequently, numerous approaches have been proposed to address these issues. In this paper, we conduct a large-scale empirical study on 16 state-of-the-art DL model fixing approaches, spanning model-level, layer-level, and neuron-level categories, to comprehensively evaluate their performance. We assess not only their fixing effectiveness (their primary purpose) but also their impact on other critical properties, such as robustness, fairness, and backward compatibility. To ensure comprehensive and fair evaluation, we employ a diverse set of datasets, model architectures, and application domains within a uniform experimental setup for experimentation. We summarize several key findings with implications for both industry and academia. For example, model-level approaches demonstrate superior fixing effectiveness compared to others. No single approach can achieve the best fixing performance while improving accuracy and maintaining all other properties. Thus, academia should prioritize research on mitigating these side effects. These insights highlight promising directions for future exploration in this field.

模型修复深度学习可靠性实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。