arXiv:2411.00259cs.LG2024-11NeurIPS被引 3

用球面能量优化CKA,提升贝叶斯深度学习的多样性与稳定性。

Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA

  • 在CKA基础上引入球面能量,增强网络间差异性度量
  • 在异常检测任务中显著提升不确定性量化性能
  • 适合需要高可靠性的模型集成与鲁棒推理场景

基于粒子的贝叶斯深度学习通常需要相似性度量来比较网络,但传统方法缺乏排列不变性且不适用于网络对比。中心核对齐(CKA)虽可用于深度网络比较,却未被用作贝叶斯学习中的优化目标。本文探索将CKA作为目标,在生成多样化的集成模型和输出网络后验的超网络中应用。注意到CKA将核映射到单位超球面,直接优化会导致网络过于相似时梯度消失。为此,我们提出在CKA核上引入球面能量(HE)机制,以缓解该问题并提升训练稳定性。此外,利用基于CKA的特征核,推导出应用于合成异常样本的特征排斥项。在多样化集成与超网络上的实验表明,本方法在合成与真实异常检测任务中均显著优于基线,在不确定性量化方面表现更优。

原文摘要 · Abstract (English)

Particle-based Bayesian deep learning often requires a similarity metric to compare two networks. However, naive similarity metrics lack permutation invariance and are inappropriate for comparing networks. Centered Kernel Alignment (CKA) on feature kernels has been proposed to compare deep networks but has not been used as an optimization objective in Bayesian deep learning. In this paper, we explore the use of CKA in Bayesian deep learning to generate diverse ensembles and hypernetworks that output a network posterior. Noting that CKA projects kernels onto a unit hypersphere and that directly optimizing the CKA objective leads to diminishing gradients when two networks are very similar. We propose adopting the approach of hyperspherical energy (HE) on top of CKA kernels to address this drawback and improve training stability. Additionally, by leveraging CKA-based feature kernels, we derive feature repulsive terms applied to synthetically generated outlier examples. Experiments on both diverse ensembles and hypernetworks show that our approach significantly outperforms baselines in terms of uncertainty quantification in both synthetic and realistic outlier detection tasks.

贝叶斯深度学习多样性增强不确定性量化CKA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。