arXiv:2607.20993cs.CVcs.AI2026-07

发现冷冻医学视觉模型中关键病灶由约10个通道稀疏编码,可精准定位并提升报告生成效率。

Sparse Concept Channels in Frozen 3D CT Vision Encoders

论文配图:Sparse Concept Channels in Frozen 3D CT Vision Encoders
图 1 · 摘自论文原文
  • 通过无训练探针识别出每种影像诊断对应约10个稀疏激活的编码通道。
  • 关闭特定通道使对应诊断得分崩溃,而其他标签保持稳定,证明编码可解释性。
  • 方法通用性强,跨模型复现且推理速度提升22倍,适合临床部署与可解释性研究。

大型视觉语言模型在三维医学图像解读中日益主导,但我们很少了解哪些内部单元编码了临床发现,以及这些信息存在于表征的何处。本研究以3D胸部视觉语言模型Pillar-0为例,探测其冻结的视觉嵌入。结果表明:(i) 每种放射学发现由约10个稀疏的视觉编码器通道编码,性能接近全特征分类,远超零样本文本提示;(ii) 关闭与某一发现相关的通道后,该发现的评分急剧下降,而无关标签保持稳定;(iii) 相同的稀疏探针在结构不同的3D腹部VLM Merlin上成功复现,表明这是冻结医学编码器的普遍特性。所提出的无训练概念通道探针(CCP)方法结合语料库生成的报告模板,在临床有效性和自然语言生成指标上优于已发表的CT-CHAT:F1达0.549(原为0.184),BLEU达0.483(原为0.373),同时实现22倍更低延迟。结果清晰、可复现地揭示了冻结医学编码器如何表征临床发现,具有跨模型直接应用价值。

原文摘要 · Abstract (English)

Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know <i>which</i> internal units encode clinical findings or <i>where</i> that information lives in the representation. We first study this on a 3D chest vision-language model (Pillar-0) by probing its frozen vision embeddings. We show that (i) each radiological finding is encoded by a <i>sparse</i> set of ~10 vision-encoder channels that match full-feature classification performance and far exceed a zero-shot text prompting; (ii) turning off the channels tied to one finding, that finding's score collapses while unrelated labels stay stable; and (iii) the same sparse probe <i>replicates</i> on an architecturally unrelated 3D abdominal VLM (Merlin) suggesting a general property of frozen medical encoders. Our training-free concept channel probe (CCP) method, paired with a corpus-derived report template, outperforms published CT-CHAT on clinical efficacy and NLG metrics (F1 0.549 vs. 0.184; BLEU 0.483 vs. 0.373) at 22x lower latency. Our results provide a clear, reproducible characterization of how frozen medical encoders represent findings, demonstrating direct applicability across models.

医学影像可解释性稀疏编码视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。