arXiv:2609.05937cs.CV2026-09

用图文对预训练让医疗影像模型聚焦病灶,提升诊断准确率和可解释性。

Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

论文配图:Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics
图 1 · 摘自论文原文
  • 用图文配对实现跨模态对齐,引导模型关注真实病变区域。
  • 在多个临床分类任务中,准确率显著优于传统分类预训练方法。
  • 适合需要高可解释性的医学影像分析场景,如辅助诊断系统。

视觉变换器(ViTs)在医疗图像分析中潜力巨大,但传统的全局图像分类预训练易导致空间坍缩,使模型依赖背景线索而非定位关键病灶。为克服此问题并使视觉证据与精准医学语义对齐,我们系统评估了三种结构化监督方式:基于图自监督的拓扑先验、像素级分割约束,以及基于图像-文本对的跨模态语义对齐。实验表明,尽管三者均缓解了全局池化瓶颈并引导注意力至前景区域,但图像-文本对齐表现最优。该方法通过嵌入高维诊断逻辑,不仅锚定注意力于精确视觉证据,还支持深层抽象推理。大量实验显示,这种语义丰富化的预训练显著提升了特征表示能力。微调至下游临床分类任务时,模型准确率更高,且注意力图高度聚焦真实病理特征,远超标准分类基线。

原文摘要 · Abstract (English)

Vision Transformers (ViTs) have shown immense potential in medical image analysis. However, standard pre-training via global image classification suffers from spatial collapse, where models rely heavily on background shortcuts rather than localising critical foreground lesions. To overcome this limitation and align visual evidence with precise medical semantics, we systematically investigate alternative pre-training paradigms.Specifically, we evaluate three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs. Notably, our empirical analysis reveals that while all three forms of structured supervision successfully alleviate the global pooling bottleneck and steer visual attention towards foreground regions, image-text alignment achieves the most superior performance. By embedding high-dimensional diagnostic logic, the cross-modal approach not only anchors attention on precise visual evidence but also enables profound abstract reasoning. Extensive experiments demonstrate that this semantically enriched pre-training fundamentally enhances the model's feature representation. Consequently, when fine-tuned for downstream clinical classification tasks, our models achieve superior accuracy and yield highly interpretable attention maps focused on true pathological features, vastly outperforming vanilla classification baselines.

视觉变换器医疗影像图文对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。