arXiv:2510.11296cs.CVcs.LG2025-10NeurIPS被引 2

提出ΔEnergy新指标,提升视觉语言模型在分布外数据上的检测与泛化能力。

$Δ\mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization

  • 基于模态对齐时的能量变化设计新OOD评分ΔEnergy
  • 在多个基准上实现10%~25%的AUROC提升
  • 适合关注模型鲁棒性与开集识别的研究者

近期视觉语言模型(VLMs)在快速下游适配方面表现优异。但在真实任务中,模型会同时遇到分布内(ID)和分布外(OOD)数据,后者包含协变量偏移(如图像风格变化)和语义偏移(如测试时未见类别)。这凸显了提升模型对协变量偏移的泛化能力,以及有效检测语义偏移类别的必要性。受闭集数据中模态重新对齐时能量显著变化的启发(具体表现为最大余弦相似度降至低值),本文提出新型OOD得分ΔEnergy,显著优于原始能量基方法,且更可靠。此外,通过下界最大化ΔEnergy(称为EBM),可同步提升对协变量偏移的泛化性能,理论证明其能生成域一致的海森矩阵,是泛化性的强指标。基于此,我们构建统一微调框架,在挑战性OOD检测与泛化基准上验证效果,性能超越现有方法10%至25%(以AUROC计)。

原文摘要 · Abstract (English)

Recent approaches for vision-language models (VLMs) have shown remarkable success in achieving fast downstream adaptation. When applied to real-world downstream tasks, VLMs inevitably encounter both the in-distribution (ID) data and out-of-distribution (OOD) data. The OOD datasets often include both covariate shifts (e.g., known classes with changes in image styles) and semantic shifts (e.g., test-time unseen classes). This highlights the importance of improving VLMs' generalization ability to covariate-shifted OOD data, while effectively detecting open-set semantic-shifted OOD classes. In this paper, inspired by the substantial energy change observed in closed-set data when re-aligning vision-language modalities (specifically by directly reducing the maximum cosine similarity to a low value), we introduce a novel OOD score, named ΔEnergy. ΔEnergy significantly outperforms the vanilla energy-based OOD score and provides a more reliable approach for OOD detection. Furthermore, ΔEnergy can simultaneously improve OOD generalization under covariate shifts, which is achieved by lower-bound maximization for ΔEnergy (termed EBM). EBM is theoretically proven to not only enhance OOD detection but also yields a domain-consistent Hessian, which serves as a strong indicator for OOD generalization. Based on this finding, we developed a unified fine-tuning framework that allows for improving VLMs' robustness in both OOD generalization and OOD detection. Extensive experiments on challenging OOD detection and generalization benchmarks demonstrate the superiority of our method, outperforming recent approaches by 10% to 25% in AUROC.

视觉语言模型分布外检测模型鲁棒性能量方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。