arXiv:2609.08524cs.CV2026-09

通过多分辨率熵估计,选出医学图像中最佳视觉语言模型中间层,提升零样本分布外检测精度。

Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation

论文配图:Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation
图 1 · 摘自论文原文
  • 采用多分辨率熵估计,克服单分辨率下分箱敏感问题。
  • 在MIDOG和OASIS数据集上,性能较当前最优提升最高达19.3% AUROC。
  • 适用于医疗影像跨机构、跨协议的零样本分布外检测场景。

分布外(OOD)检测对安全部署医学AI系统至关重要,因机构、采集协议及患者群体差异常引发领域偏移。视觉语言模型(VLM)可通过将图像嵌入与语言对齐的潜在空间,利用跨模态相似性实现零样本OOD检测,作为非参数置信度信号识别分布内样本。然而现有方法几乎仅依赖最终层嵌入,隐含假设最深层表征始终最优。我们首次发现该假设在医学影像中不成立:中间层提供互补的OOD信号,最优表征深度取决于图像模态。先前工作通过归一化直方图熵最小化选择层组合,但单分辨率熵估计对分箱高度敏感,导致AUROC波动高达19.3%。为此,我们提出多分辨率熵估计策略,聚合多尺度离散化下的直方图统计,实现鲁棒稳定的中间层选择。在涵盖不同成像模态、多样偏移类型与多种VLM主干的两个医学OOD基准(MIDOG与OASIS)上,本方法持续优于现有最优方案,提供轻量且稳定的零样本OOD检测解决方案。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection is crucial for safe deployment of medical AI systems, where domain shifts arise across institutions, acquisition protocols, and patient populations. VLMs enable zero-shot OOD detection by embedding images into a language-aligned latent space, where cross-modal similarity serves as a non-parametric confidence signal for identifying in-distribution samples. Yet existing methods rely almost exclusively on final-layer embeddings, implicitly assuming that the deepest representations are universally optimal. We first show that this assumption does not hold in medical imaging: intermediate layers provide complementary OOD signals, and the optimal representational depth depends on the respective image modality. While prior work selects layer combinations via entropy minimization of normalized histograms, we demonstrate that single-resolution entropy estimation is highly sensitive to binning choices, leading to performance variations of up to 19.3% AUROC. To address this instability, we propose a multi-resolution entropy estimation strategy that aggregates histogram statistics across multiple discretization scales, enabling robust and stable intermediate-layer selection. Across two medical OOD benchmarks, namely MIDOG and OASIS, covering distinct imaging modalities, diverse shift types, and different VLM backbones, our method consistently outperforms state-of-the-art approaches, offering a lightweight and stable solution for zero-shot OOD detection.

医学AIOOD检测视觉语言模型多分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。