arXiv:2604.13021cs.CVcs.AI2026-04

研究医学影像中视觉语言模型的表示几何,发现不同聚合方式影响诊断与检索表现。

Representation geometry shapes task performance in vision-language modeling for CT enterography

  • 用均值池化和注意力池化分别提升分类与跨模态检索性能。
  • 多窗RGB编码优于增加扫描视角,且额外切面降低分类准确率。
  • 检索增强生成显著提升报告生成质量,适合临床辅助决策系统。

计算机断层扫描(CT)小肠造影是评估炎症性肠病的主要影像手段,但其自动化分析的最佳表征选择尚不明确。本研究首次开展腹部CT小肠造影的视觉语言迁移学习,发现:首先,对切片嵌入采用均值池化可获得59.2%的三分类准确率,而注意力池化在跨模态检索上表现更优(文本到图像MRR为0.235),该趋势在所有LoRA配置下均成立,表明两种聚合器强调不同表征特性;其次,单切片组织对比度比广域空间覆盖更重要:多窗RGB编码(将互补的亨氏单位窗口映射至RGB通道)优于通过多平面采样扩大空间覆盖的所有策略,且添加冠状面与矢状面反而降低分类性能。对于报告生成,无检索上下文微调仅达70.4%的严重程度准确率,接近随机水平(71%);而检索增强生成(RAG)在所有配置下均提升7–14个百分点,有序误差从0.98降至0.80–0.89。基于三教师伪标签框架实现无需专家标注的全量比较。这些发现为该未充分探索的模态提供了首个基线,并为构建体积医学影像的视觉语言系统提供实用指导。

原文摘要 · Abstract (English)

Computed tomography (CT) enterography is a primary imaging modality for assessing inflammatory bowel disease (IBD), yet the representational choices that best support automated analysis of this modality are unknown. We present the first study of vision-language transfer learning on abdominal CT enterography and identify two main findings. First, mean pooling of slice embeddings gives better categorical disease assessment (59.2\% three-class accuracy), whereas attention pooling gives better cross-modal retrieval (0.235 text-to-image MRR). This pattern holds across all LoRA configurations tested and suggests that the two aggregators emphasize different properties of the learned representation. Second, per-slice tissue contrast matters more than broader spatial coverage: multi-window RGB encoding, which maps complementary Hounsfield Unit windows to RGB channels, outperforms all strategies that increase spatial coverage through multiplanar sampling, and in this setting adding coronal and sagittal views reduces classification performance. For report generation, fine-tuning without retrieval context yields within-1 severity accuracy at the prevalence-matched chance level (70.4\% vs.\ 71\% random), suggesting little learned ordering beyond the class distribution. Retrieval-augmented generation (RAG) improves this across all configurations, scoring 7--14 percentage points above the chance baseline and improving ordinal MAE from 0.98 to 0.80--0.89. A three-teacher pseudolabel framework enables all comparisons without expert annotations. Together, these findings provide the first baselines for this underexplored modality and offer practical guidance for building vision-language systems for volumetric medical imaging.

医学影像视觉语言表征学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。