arXiv:2512.11141cs.CV2025-12

用分项文本监督训练可解释的视觉模型,提升医学影像理解能力

Learning complete and explainable visual representations from itemized text supervision

  • 通过交叉注意力生成分项文本条件下的视觉嵌入
  • 在脑部MRI等4个领域零样本性能显著优于基线
  • 适合需要细粒度可解释性的医疗与遥感图像分析

使用语言监督训练视觉模型可获得通用且可迁移的表示。然而,许多视觉领域(如医学影像和遥感)包含分项文本标注:单张图像中多个语义独立的发现由多个文本项描述。此类监督不同于标准多标题监督(标题冗余或高度重叠)。本文提出ItemizedCLIP框架,从分项文本监督中学习完整且可解释的视觉表示。该框架采用交叉注意力模块生成文本项条件下的视觉嵌入,并设计一组目标函数,联合强制实现项独立性(不同文本项对应不同区域)和表示完整性(覆盖所有文本项)。在四个天然存在分项文本标注的领域(脑部MRI、头颅CT、胸部CT、遥感)及一个合成分项数据集上,ItemizedCLIP在零样本性能和细粒度可解释性方面均显著优于基线。所得表示具有语义基础、项可区分、完整且可视可解释。代码已开源。

原文摘要 · Abstract (English)

Training vision models with language supervision enables general and transferable representations. However, many visual domains, especially non-object-centric domains such as medical imaging and remote sensing, contain itemized text annotations: multiple text items describing distinct and semantically independent findings within a single image. Such supervision differs from standard multi-caption supervision, where captions are redundant or highly overlapping. Here, we introduce ItemizedCLIP, a framework for learning complete and explainable visual representations from itemized text supervision. ItemizedCLIP employs a cross-attention module to produce text item-conditioned visual embeddings and a set of tailored objectives that jointly enforce item independence (distinct regions for distinct items) and representation completeness (coverage of all items). Across four domains with naturally itemized text supervision (brain MRI, head CT, chest CT, remote sensing) and one additional synthetically itemized dataset, ItemizedCLIP achieves substantial improvements in zero-shot performance and fine-grained interpretability over baselines. The resulting ItemizedCLIP representations are semantically grounded, item-differentiable, complete, and visually interpretable. Our code is available at https://github.com/MLNeurosurg/ItemizedCLIP.

视觉表示可解释性医学影像文本监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。