提出方向性尖锐度,高效可靠评估模型泛化能力。
Certification of Machine Learning Models via Directional Sharpness
- 引入方向性尖锐度,通过训练数据验证模型泛化性能。
- 相比传统指标,与泛化能力相关性更强且更稳定。
- 支持审计与零知识证明,保护数据隐私的同时认证模型。
在机器学习中,模型认证是确保模型可信性和质量的重要方法。模型质量主要取决于其泛化能力,即在训练数据之外的表现。然而,泛化能力无法直接认证,因其依赖未知数据且不可直接测量。测试准确率等代理指标在训练过程受扰动时可能误导,而现有尖锐度指标虽与泛化有经验关联,但计算成本高,且在训练偏离标准流程时不可靠。本文提出方向性尖锐度,一种可在训练偏差下仍高效、可靠指示泛化能力的度量。我们提供了实证和分析证据表明:(1) 方向性尖锐度与泛化相关性优于现有指标;(2) 更能可靠识别泛化能力差的模型。此外,该度量在模型审计场景中可高效计算,并可通过零知识证明实现无需披露训练数据的质量认证。
原文摘要 · Abstract (English)
In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality. A model's quality is largely determined by its ability to generalize, i.e., to perform well on data beyond what it was trained on. It is not possible to certify generalization directly, however, as it depends on unknown data and is not directly measurable. Proxies such as test accuracy can be misleading when the training process is perturbed (intentionally or accidentally), and metrics such as sharpness -- which has an empirically supported link to generalization -- are computationally expensive and can also serve as unreliable signals when training deviates from a prescribed procedure. In this work, we propose directional sharpness, a metric designed to efficiently and reliably indicate generalization despite potential training deviations. We provide empirical and analytical evidence that directional sharpness (1) correlates more strongly with generalization than existing metrics and (2) identifies models with poor generalization more reliably than existing metrics. Furthermore, directional sharpness is efficiently computable in model auditing settings, where the verifier has access to training data, and via zero-knowledge proofs that certify quality without revealing training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。