arXiv:2605.24771cs.CVcs.AI2026-05

为医学影像弱监督设定可落地的标签质量判断标准

From Theory to Decision Rule: Calibrating the Noisy-Label Crossover for Vision-Language Model Weak Supervision Across Three Medical-Imaging Benchmarks

  • 基于三组医学影像数据验证理论交叉点,给出具体阈值
  • 超过阈值后标签使模型性能下降最高达0.10 AUC
  • 仅需10-20个真实标签即可判断是否应使用弱标签

经典噪声标签理论预测,在弱监督下下游性能受标注者准确率上限制约,存在明确的交叉点:一旦黄金训练分类器达到标注者水平,弱标签便不再有益反而有害。该理论尚缺实际基准校准,无法转化为现代基础模型标注者的实例级判断。本文针对BiomedCLIP生成的弱标签,在PCAM、ISIC、NIH-CXR三个医学影像基准上完成校准,覆盖六种下游架构(参数量相差11倍)。理论预测的交叉点在PCAM约为100,ISIC为20-50,NIH-CXR为250-500;超过该阈值时,弱标签使AUC下降最多-0.10。该交叉点对五种预训练架构中的四种保持不变,且同一家族DenseNet的参数量差异(2.5倍)实验支持标注者是主要约束因素。由此构建可操作决策规则:仅需10-20个真实标签,比较黄金集上的模型性能与视觉语言模型在该集上的准确率即可。在NIH-CXR上结构化与随机噪声对比表明,仅考虑噪声率的边界形式不完整,并提出标签空间投影作为可测试的改进方向。

原文摘要 · Abstract (English)

Classical noisy-label theory predicts that downstream performance under weak supervision is bounded above by the labeler's accuracy, implying a sharp crossover: once a gold-trained classifier matches the labeler, weak labels stop helping and start hurting. The prediction is theoretical; what is missing is a benchmark calibration that turns it into an instance-level statement for modern foundation-model labelers. We provide such a calibration for BiomedCLIP-generated weak labels on three medical-imaging benchmarks (PCAM, ISIC, NIH-CXR) and six downstream architectures spanning an 11x parameter range. The crossover predicted by theory appears at ng~100 on PCAM, 20-50 on ISIC, and 250-500 on NIH-CXR; weak labels above the crossover degrade AUC by up to -0.10. The location is architecture-invariant for four of five pretrained architectures, and a within-family DenseNet sweep (2.5x parameters, identical pretraining) supports the view that the labeler, not the student, is the dominant constraint. The calibration in turn produces a decision rule operable from 10-20 gold labels: compare gold-only AUC to VLM accuracy on the user's gold set. A structured-vs-random noise sign flip on NIH-CXR shows that the rate-only formulation of the bound is incomplete and identifies a concrete refinement (label-space projection) that future benchmarks can be designed to test.

弱监督医学影像标签校准VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。