让CT模型同时学习11个器官的专属特征,提升疾病诊断精度。
OrganLens: Organ-Specific Representation Learning for CT Foundation Models

- 用器官身份控制共享编码器,自监督学习器官特异性表征。
- 心肺等器官的诊断性能显著提升,如心脏病变识别准确率提高至0.953。
- 无需分割掩码即可生成器官级表示,适合临床研究与多任务分析。
CT检查包含多个器官,但许多生物医学问题聚焦于特定器官的异常、预后或纵向变化,需要在同一体积内为每个器官建立独立表征。现有CT基础模型通常生成单一整体表征,而近期解剖感知方法要么对分离的器官体积编码,要么显式将图像拆分为器官令牌组。前者可能丢失关键邻近上下文,后者未在特征形成前以目标器官条件化共享编码器。本文提出OrganLens,通过自监督实现器官特异性表征学习:器官身份条件化共享CT编码器,结合器官特异性蒸馏与解剖掩码监督,经解剖加权池化生成器官特异性表征。推理时,共享模型无需外部分割掩码即生成11个器官的表征。在CT-RATE、RAD-ChestCT、INSPECT和NLST等多个数据集上评估,相比预训练的DINOv2,心脏表征使心室扩大识别的AUROC从0.910提升至0.953,肺部表征使NLST肺癌死亡风险预测的Harrell C-index提高14.2%。全局表征在文本到图像与图像到文本检索任务中分别达到33.09%与32.04%的Recall@10。在器官相关任务中,解剖匹配的表征提供更强任务信号,而全局表征仍具广泛适用性。OrganLens为基于共享编码器的器官特异性CT表征学习提供可扩展方案,也为跨队列、跨临床终点的器官特异性疾病研究提供可复用框架。
原文摘要 · Abstract (English)
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ within the same CT volume. Existing CT foundation models commonly produce a single volume-level representation, while recent anatomy-aware methods either encode pre-separated organ volumes or explicitly disentangle images into organ token groups. The former may remove clinically relevant surrounding context, while the latter does not condition a shared encoder on a selected organ before its features are formed. We introduce OrganLens for organ-specific representation learning through self-supervision. An organ identity conditions a shared CT encoder, while organ-specific distillation and anatomy-mask supervision shape features for anatomy-weighted pooling into organ-specific representations. At inference, the shared model produces 11 organ-specific representations without external segmentation masks. We evaluate OrganLens on CT-RATE, RAD-ChestCT, INSPECT, and NLST across diverse acquisitions and downstream evaluations. Relative to CT-pretrained DINOv2, heart representations raise CT-RATE cardiomegaly AUROC from 0.910 to 0.953, while lung representations improve the Harrell C-index for NLST lung-cancer mortality by 14.2\%. The global representation reaches INSPECT Recall@10 of 33.09\% and 32.04\% for text-to-image and image-to-text retrieval, respectively. Across organ-related tasks, anatomically matched representations provide stronger task-relevant signal, while the global representation retains broad utility. OrganLens offers a scalable approach to organ-specific CT representation learning with a shared encoder. More broadly, it provides the medical research community with a reusable framework for studying organ-specific disease across cohorts and clinical endpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。