提出新型证据深度学习模型,解决不确定性建模中的梯度消失问题。
Generalized Regularized Evidential Deep Learning Models: Theory and Comprehensive Evaluation
- 设计通用激活函数族与正则化项,突破传统证据非负约束。
- 在多个数据集上验证新模型有效避免低证据区的梯度消失。
- 适合需要可靠不确定估计的高风险场景如医疗诊断、自动驾驶。
基于主观逻辑的证据深度学习(EDL)模型为确定性神经网络提供了系统且计算高效的不确定性感知方法,可通过学习到的证据量化细粒度不确定性。然而,主观逻辑框架要求证据非负,需使用特定激活函数,其几何特性可能导致激活依赖的学习冻结现象:在低证据区域样本的梯度极小,阻碍学习。本文理论刻画该行为,并分析不同证据激活函数对学习动态的影响。基于此分析,我们设计了一类通用激活函数及其对应的证据正则化项,提供跨激活区域一致证据更新的新路径。在四个基准分类任务(MNIST、CIFAR-10、CIFAR-100、Tiny-ImageNet)、两个少样本分类任务及盲人脸恢复任务上的大量实验,验证了所提理论并证明了新模型的有效性。
原文摘要 · Abstract (English)
Evidential deep learning (EDL) models, based on Subjective Logic, introduce a principled and computationally efficient way to make deterministic neural networks uncertainty-aware. The resulting evidential models can quantify fine-grained uncertainty using learned evidence. However, the Subjective-Logic framework constrains evidence to be non-negative, requiring specific activation functions whose geometric properties can induce activation-dependent learning-freeze behavior: a regime where gradients become extremely small for samples mapped into low-evidence regions. We theoretically characterize this behavior and analyze how different evidential activations influence learning dynamics. Building on this analysis, we design a general family of activation functions and corresponding evidential regularizers that provide an alternative pathway for consistent evidence updates across activation regimes. Extensive experiments on four benchmark classification problems (MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet), two few-shot classification problems, and blind face restoration problem empirically validate the developed theory and demonstrate the effectiveness of the proposed generalized regularized evidential models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。