通过词元级不确定性量化,评估视觉语言模型在病理图像分析中的可信度。
Logit-Level Uncertainty Quantification in Vision-Language Models for Histopathology Image Analysis
- 基于温度调节的词元输出,量化模型不确定性
- 不同模型表现差异大,通用模型易出现突变不确定性
- 专用模型稳定性高,适合临床可信诊断场景
视觉语言模型(VLMs)在教育、交通、医疗、能源、金融、法律和零售等多个领域展现出卓越性能。然而,其在医疗应用中使用时面临大规模医疗数据敏感性及模型可信度(可靠性、透明性、安全性)的关键挑战。本研究提出一种针对病理图像分析的词元级不确定性量化(UQ)框架,以应对上述问题。通过温度控制输出词元的指标评估三种VLMs的UQ表现。结果显示,模型间存在显著差异:通用模型如VILA-M3-8B与LLaVA-Med v1.5表现出高随机敏感性(平均余弦相似性CS <0.71, <0.84;JS散度<0.57, <0.38;KL散度<0.55, <0.35),温度变化影响接近最大(Δ_T ≈1.00),且复杂诊断提示下出现突变不确定性;而专用于病理的PRISM模型则保持近确定性行为(平均CS >0.90,JS <0.10,KL <0.09),对温度变化响应极小,适用于各类提示。结果强调了词元级不确定性量化在评估病理领域VLM可信度中的重要性。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) with their multimodal capabilities have demonstrated remarkable success in almost all domains, including education, transportation, healthcare, energy, finance, law, and retail. Nevertheless, the utilization of VLMs in healthcare applications raises crucial concerns due to the sensitivity of large-scale medical data and the trustworthiness of these models (reliability, transparency, and security). This study proposes a logit-level uncertainty quantification (UQ) framework for histopathology image analysis using VLMs to deal with these concerns. UQ is evaluated for three VLMs using metrics derived from temperature-controlled output logits. The proposed framework demonstrates a critical separation in uncertainty behavior. While VLMs show high stochastic sensitivity (cosine similarity (CS) $<0.71$ and $<0.84$, Jensen-Shannon divergence (JS) $<0.57$ and $<0.38$, and Kullback-Leibler divergence (KL) $<0.55$ and $<0.35$, respectively for mean values of VILA-M3-8B and LLaVA-Med v1.5), near-maximal temperature impacts ($Δ_T \approx 1.00$), and displaying abrupt uncertainty transitions, particularly for complex diagnostic prompts. In contrast, the pathology-specific PRISM model maintains near-deterministic behavior (mean CS $>0.90$, JS $<0.10$, KL $<0.09$) and significantly minimal temperature effects across all prompt complexities. These findings emphasize the importance of logit-level uncertainty quantification to evaluate trustworthiness in histopathology applications utilizing VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。