arXiv:2608.05960cs.CVcs.AI2026-08

评测10个3D CT模型发现:小而低对比度病灶难检测,性能关键在可辨识度而非模型大小。

Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models

论文配图:Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
图 1 · 摘自论文原文
  • 用k近邻、零样本提示等方法评估10个冻结特征的CT模型在肺部扫描中的表现。
  • 高对比或范围广的异常(如积液、支架)能被稳定识别,小病灶仍普遍漏检。
  • 模型效果主要受病灶与周围组织对比度和空间范围影响,非架构或规模决定。

常规CT解读需全面覆盖扫描体积以发现意外发现。3D CT基础模型可通过提供解剖与病理的通用表征辅助此过程。为评估其诊断广度,我们在三个胸腔CT队列上对十个冻结特征的编码器进行评测,包括一个未见过的内部临床数据集,采用k-最近邻、零样本提示和线性探测方法。结果显示不存在统一的最优模型,排名随评估上下文显著波动。结合细粒度图像标记与视觉-语言对齐的模型通常表现最佳,但轻量级监督编码器同样具有竞争力,表明显式标注可替代规模优势。关键发现是:性能主因并非模型架构,而是物理瓶颈——病灶可检测性与其相对于周围组织的对比度及空间范围成正比。通过器官内受控对比实验,我们证实广泛或高对比异常(如器械、积液)可被可靠恢复;相反,小且低对比的局灶性病变在所有模型中仍是持续挑战。这归因于全局池化嵌入的固有局限,表明准确表征小而低对比结构需区域或病灶级预训练。

原文摘要 · Abstract (English)

Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across three cohorts of thoracic CT scans, including an unseen internal clinical dataset, using $k$-nearest neighbors, zero-shot prompting, and linear probing. We find no universal state-of-the-art, with rankings fluctuating significantly depending on the evaluation context. While models combining fine-grained image tokenization with vision-language alignment generally perform best, a lightweight supervised encoder remains highly competitive, demonstrating that explicit labels can effectively substitute for scale. Crucially, rather than model architecture, we observe that the primary determinant of performance is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent. Through controlled within-organ comparisons, we empirically demonstrate that widespread or high-contrast abnormalities, such as devices and effusions, are reliably recovered. Conversely, small, low-contrast focal lesions remain a persistent challenge across all evaluated encoders. We attribute this to the inherent limitations of globally pooled embeddings, suggesting that accurately representing small, low-contrast structures will require region- or lesion-level pretraining.

3D CT基础模型病灶检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。