arXiv:2409.01498cs.LG2024-09被引 4

提出可量化模型泛化能力的新指标,用于评测深度网络表现。

A practical generalization metric for deep networks benchmarking

  • 基于准确率与未知数据多样性构建泛化评估指标。
  • 实测发现多数理论估计与实际表现无相关性。
  • 适合模型对比与理论验证,推动泛化研究落地。

当前对深度学习模型泛化误差的理论边界估计持续受到关注,同时实践性评估指标的需求也在增长。这一需求不仅出于实际应用考量,也对理论研究至关重要,因理论成果需通过实践验证。然而,目前缺乏对不同深度网络泛化能力的系统性基准测试,也缺少对理论估计的有效验证手段。本文提出一种实用的泛化度量方法,并构建新的测试平台以验证理论估计。研究发现,深度网络在分类任务中的泛化能力取决于分类准确率和未见数据的多样性。所提指标可同时量化模型精度与数据多样性,提供直观且量化的评估方式,形成权衡点。我们使用该测试平台将新指标与现有理论估计进行比较,结果令人失望:多数理论估计与本研究所得实际测量值无相关性。这一发现揭示了现有理论估计的局限性,也为未来研究提供了新方向。

原文摘要 · Abstract (English)

There is an ongoing and dedicated effort to estimate bounds on the generalization error of deep learning models, coupled with an increasing interest with practical metrics that can be used to experimentally evaluate a model's ability to generalize. This interest is not only driven by practical considerations but is also vital for theoretical research, as theoretical estimations require practical validation. However, there is currently a lack of research on benchmarking the generalization capacity of various deep networks and verifying these theoretical estimations. This paper aims to introduce a practical generalization metric for benchmarking different deep networks and proposes a novel testbed for the verification of theoretical estimations. Our findings indicate that a deep network's generalization capacity in classification tasks is contingent upon both classification accuracy and the diversity of unseen data. The proposed metric system is capable of quantifying the accuracy of deep learning models and the diversity of data, providing an intuitive and quantitative evaluation method, a trade-off point. Furthermore, we compare our practical metric with existing generalization theoretical estimations using our benchmarking testbed. It is discouraging to note that most of the available generalization estimations do not correlate with the practical measurements obtained using our proposed practical metric. On the other hand, this finding is significant as it exposes the shortcomings of theoretical estimations and inspires new exploration.

泛化能力深度学习评估指标模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。