通过优化混淆矩阵最大特征值,提升平滑分类器最差类别的鲁棒性。
Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers
- 基于最大特征值设计正则化方法,改善平滑分类器的最差类别表现。
- 理论证明最大特征值直接影响最差类别错误率,实验证明其有效性。
- 适合关注模型公平性和鲁棒性均衡的研究者与应用开发者。
近期研究揭示了深度神经网络中一个关键挑战——‘鲁棒不公平’现象,即模型在不同类别间的鲁棒准确率存在显著差异。尽管已有工作尝试缓解对抗鲁棒性中的不公,但平滑分类器的最差类别认证鲁棒性研究仍属空白。本文通过构建平滑分类器最差类别误差的PAC-Bayesian界,理论分析表明:平滑混淆矩阵的最大特征值从根本上影响最差类别误差。基于此,我们提出一种正则化方法,优化平滑混淆矩阵的最大特征值,以提升平滑分类器的最差类别准确率,并进一步增强其最差类别认证鲁棒性。我们在多个数据集和模型架构上进行了广泛实验,验证了该方法的有效性。
原文摘要 · Abstract (English)
Recent studies have identified a critical challenge in deep neural networks (DNNs) known as ``robust fairness", where models exhibit significant disparities in robust accuracy across different classes. While prior work has attempted to address this issue in adversarial robustness, the study of worst-class certified robustness for smoothed classifiers remains unexplored. Our work bridges this gap by developing a PAC-Bayesian bound for the worst-class error of smoothed classifiers. Through theoretical analysis, we demonstrate that the largest eigenvalue of the smoothed confusion matrix fundamentally influences the worst-class error of smoothed classifiers. Based on this insight, we introduce a regularization method that optimizes the largest eigenvalue of smoothed confusion matrix to enhance worst-class accuracy of the smoothed classifier and further improve its worst-class certified robustness. We provide extensive experimental validation across multiple datasets and model architectures to demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。