arXiv:2607.28696cs.LGcs.CV2026-07

提出新方法提升医疗视觉语言模型在临床数据漂移下的疾病类别覆盖率。

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

论文配图:Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift
图 1 · 摘自论文原文
  • 基于交叉验证识别高风险类别,结合局部与尾部阈值保护
  • 在8个设置下实现0.95以上边际和最差类别覆盖率
  • 适合关注医疗AI可靠性与公平性的研究者与临床应用

医疗视觉语言模型在临床数据漂移后虽保持总体覆盖率,但个别疾病类别严重覆盖不足。该问题随采集协议与模型结构变化,无法通过原始发病率判断。现有局部与尾部感知校准方法分别应对测试邻域与频次尾部,但未能建模类别级遗漏。本文提出CALCoDe,一种针对冻结医疗VLM的后处理可靠性层。通过交叉验证预测识别高风险类别,独立校准集估计其类别条件尾部阈值。CALCoDe将保护阈值与局部阈值取单边最大,确保集合包含所有局部规则接纳标签,额外保护仅限验证识别类别。独立支持审计机制对内点支持不足样本进行拒答。在每个保护类别的接受样本满足可交换性条件下,提供预设置信水平的有限样本覆盖率,且覆盖范围包含对应局部校准集合。在两个皮肤科数据漂移场景(HAM10000→ISIC 2019 和 HAM10000→PAD-UFES-20)及四种冻结模型(BiomedCLIP、OpenAI CLIP ViT-B/32、PubMedCLIP ViT-B/32、MedSigLIP-448)中,CALCoDe是唯一在全部八组设置下边际与最差类别接受覆盖率均达0.95的方法。在HAM10000→ISIC 2019上,其平均最差类别覆盖率0.970,优于sTACP的0.926和LCP-VLM的0.864。

原文摘要 · Abstract (English)

Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone geometry, so source prevalence does not reliably reveal the failure. Existing localized and tail-aware conformal methods respectively adapt to test neighborhoods and source-frequency tails, leaving held-out class-wise coverage failure unmodeled. We introduce Class-Tail Adaptive Localized Conformal Deferral (CALCoDe), a post-hoc reliability layer for frozen medical VLMs. Cross-fitted validation predictions identify classes at risk of undercoverage, and a disjoint calibration split estimates their class-conditional tail thresholds. CALCoDe combines each protected threshold with a localized conformal threshold using a one-sided maximum. The resulting set contains every label admitted by the localized rule, with additional protection confined to validation-identified classes. An independently calibrated support audit defers cases with insufficient inlier support. Under exchangeability among accepted examples within each protected class, CALCoDe provides finite-sample coverage at the prespecified guard level and contains the corresponding localized conformal sets; coverage on shifted external cohorts is evaluated empirically. Among standard conformal baselines and recent VLM-specific conformal methods evaluated across two dermatology shifts (HAM10000 to ISIC 2019 and HAM10000 to PAD-UFES-20) and four frozen VLM backbones (BiomedCLIP, OpenAI CLIP ViT-B/32, PubMedCLIP ViT-B/32, and MedSigLIP-448), CALCoDe is the only approach whose observed marginal and worst-class accepted coverage both reach 0.95 in all eight settings. On HAM10000 to ISIC 2019, its average worst-class accepted coverage is 0.970, compared with 0.926 for sTACP and 0.864 for LCP-VLM.

医疗AI可靠性覆盖保证视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。