arXiv:2601.00716cs.CVcs.AI2026-01被引 3

提出双路径方法,用数据分布和预测置信度联合检测病理视觉语言模型的性能退化。

Detecting Performance Degradation under Data Shift in Pathology Vision-Language Model

  • 结合输入数据分布变化与输出置信度变化,双重监测模型可靠性。
  • 在大规模肿瘤分类数据集上,双指标组合可更可靠识别性能退化。
  • 无需标签,适合临床部署后持续监控大模型表现。

视觉语言模型在医学图像分析与疾病诊断中展现出强大潜力。然而,部署后当输入数据分布发生偏移时,其性能可能下降。检测此类性能退化对临床可靠性至关重要,但对无标签数据运行的大规模预训练VLM而言仍具挑战。本研究探究了先进病理VLM在数据分布偏移下的性能退化检测。我们同时考察输入级数据偏移与输出级预测行为,以理解其在模型可靠性监控中的作用。为系统分析输入数据偏移,我们开发了DomainSAT——一个轻量级工具箱,集成主流偏移检测算法,并提供图形化界面实现直观探索。分析表明,输入数据偏移检测虽能有效识别分布变化并提供早期预警,但并不总对应实际性能下降。基于此,我们进一步研究输出端监测,提出一种无标签、基于置信度的退化指标,直接捕捉模型预测置信度的变化。结果发现该指标与性能退化密切相关,可有效补充输入偏移检测。在大规模病理肿瘤分类数据集上的实验表明,结合输入偏移检测与输出置信度指标,能更可靠地检测并解释VLM在数据偏移下的性能退化。这些发现为数字病理领域基础模型的可靠性监控提供了实用且互补的框架。

原文摘要 · Abstract (English)

Vision-Language Models have demonstrated strong potential in medical image analysis and disease diagnosis. However, after deployment, their performance may deteriorate when the input data distribution shifts from that observed during development. Detecting such performance degradation is essential for clinical reliability, yet remains challenging for large pre-trained VLMs operating without labeled data. In this study, we investigate performance degradation detection under data shift in a state-of-the-art pathology VLM. We examine both input-level data shift and output-level prediction behavior to understand their respective roles in monitoring model reliability. To facilitate systematic analysis of input data shift, we develop DomainSAT, a lightweight toolbox with a graphical interface that integrates representative shift detection algorithms and enables intuitive exploration of data shift. Our analysis shows that while input data shift detection is effective at identifying distributional changes and providing early diagnostic signals, it does not always correspond to actual performance degradation. Motivated by this observation, we further study output-based monitoring and introduce a label-free, confidence-based degradation indicator that directly captures changes in model prediction confidence. We find that this indicator exhibits a close relationship with performance degradation and serves as an effective complement to input shift detection. Experiments on a large-scale pathology dataset for tumor classification demonstrate that combining input data shift detection and output confidence-based indicators enables more reliable detection and interpretation of performance degradation in VLMs under data shift. These findings provide a practical and complementary framework for monitoring the reliability of foundation models in digital pathology.

视觉语言模型性能退化数据偏移病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。