提出可适配视觉模型的贝叶斯头,提升噪声标签下的分类鲁棒性。
Architecture-agnostic Lipschitz-constant Bayesian header and its application to resolve semantically proximal classification errors with vision transformers

- 通过谱归一化约束均值与方差,实现稳定不确定性估计。
- 在15%语义混淆标签下召回率达0.93,优于KNN方法7%以上。
- 兼容预训练模型,适合标注质量不一的高风险场景。
标签噪声是监督深度学习模型泛化能力的关键瓶颈,尤其当错误具有语义相近特性时,传统鲁棒训练方法失效。本文提出一种架构无关的利普希茨常数贝叶斯头,可集成至视觉变压器等特征提取器中,构建双利普希茨约束的贝叶斯视觉变压器(LipB-ViT)。与传统贝叶斯层不同,该方法对变分权重的均值和对数方差均施加谱归一化,增强预测不确定性的校准性并抑制噪声放大。进一步提出新指标,联合捕捉误分类率下的不确定性和置信度,并设计自适应算术平均融合策略,结合特征空间相似性与预测不确定性,识别受损标签,性能超越基于k近邻的方法超过7%,在15%语义误分类标签下召回率超0.93。尽管蒙特卡洛采样增加计算开销,但该方法具备与预训练主干网络即插即用的兼容性,且跨领域保持一致超参数,适用于标注可靠性多变的高风险应用。稳定的置信度估计构成数据集质量与标签噪声联合评估分析流程的基础,衍生出第二项新型量化指标。最后,在推理阶段系统评估了LipB-ViT在结构化(对抗性)与非结构化噪声下的表现,验证其在真实高噪声及攻击场景中的鲁棒性,并与基线方法对比。
原文摘要 · Abstract (English)
Label noise remains a critical bottleneck for the generalization of supervised deep learning models, particularly when errors are structured rather than random. Standard robust training methods often fail in the presence of such semantically proximal classification errors. This work presents an architecture-agnostic Lipschitz-constant Bayesian header that can be integrated into feature extractors such as vision transformers, yielding the bi-Lipschitz-constrained Bayesian Vision Transformer (LipB-ViT). In contrast to conventional Bayesian layers, our approach enforces spectral normalization on both the mean and log-variance of the variational weights, which promotes calibrated predictive uncertainty and mitigates noise amplification. We further propose a novel metric to jointly capture uncertainty and confidence across misclassification rates, as well as an adaptive arithmetic-mean fusion scheme that combines feature-space proximity with predictive uncertainty to detect corrupted labels outperforming the state of the art k-nearest neighbor based identification methods by more than 7% reaching a recall of more than 0.93 at 15% semantically misclassified labels. Although computational costs increase due to Monte Carlo sampling, the method offers plug-and-play compatibility with pre-trained backbones and consistent hyperparameters across domains, suggesting strong utility for high-stakes applications with variable annotation reliability. The stabilized confidence estimates serve as the foundation for an analysis pipeline that jointly assesses dataset quality and label noise, yielding a second novel metric for their combined quantification. Lastly, we systematically evaluate LipB-ViT under both structured (adversarial) and unstructured noise at inference time, demonstrating its robustness in realistic high-noise and attack scenarios. We compare its performance against baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。