用影像组学生成提示,让视觉语言模型更准判断肺结节恶性程度。
AutoRad-Lung: A Radiomic-Guided Prompting Autoregressive Vision-Language Model for Lung Nodule Malignancy Prediction
- 用手工影像组学特征生成动态提示,引导模型关注关键细节。
- 在LIDC-IDRI数据集上准确率达89.2%,优于现有CLIP类模型。
- 适合临床辅助诊断,尤其对难以区分的模糊结节有优势。
肺癌是全球癌症死亡的主要原因之一。早期诊断的关键挑战在于区分视觉特征相似、标注评分接近的不确定病例。临床上,放射科医生依赖从CT图像中提取的定量手工影像组学特征,而近年研究主要聚焦深度学习方案。近期,基于对比语言-图像预训练(CLIP)的视觉语言模型(VLM)因其可整合文本知识而受到关注。尽管CLIP-Lung模型表现良好,我们发现其存在三方面局限:(a) 依赖主观且易错的放射科医生标注属性;(b) 文本信息仅在训练阶段使用,推理时无法直接应用;(c) 视觉编码器采用随机初始化的卷积结构,忽视先验知识。为此,我们提出AutoRad-Lung,将自回归预训练的VLM与手工影像组学生成的提示相结合。该模型采用大规模自回归图像模型(AIMv2)的视觉编码器,通过多模态自回归目标预训练。由于肺部肿瘤通常小、形状不规则且与正常组织视觉相似,AutoRad-Lung能更精准捕捉像素级差异。此外,我们引入条件上下文优化机制,根据输入影像组学动态生成特定上下文提示,增强跨模态对齐能力。
原文摘要 · Abstract (English)
Lung cancer remains one of the leading causes of cancer-related mortality worldwide. A crucial challenge for early diagnosis is differentiating uncertain cases with similar visual characteristics and closely annotation scores. In clinical practice, radiologists rely on quantitative, hand-crafted Radiomic features extracted from Computed Tomography (CT) images, while recent research has primarily focused on deep learning solutions. More recently, Vision-Language Models (VLMs), particularly Contrastive Language-Image Pre-Training (CLIP)-based models, have gained attention for their ability to integrate textual knowledge into lung cancer diagnosis. While CLIP-Lung models have shown promising results, we identified the following potential limitations: (a) dependence on radiologists' annotated attributes, which are inherently subjective and error-prone, (b) use of textual information only during training, limiting direct applicability at inference, and (c) Convolutional-based vision encoder with randomly initialized weights, which disregards prior knowledge. To address these limitations, we introduce AutoRad-Lung, which couples an autoregressively pre-trained VLM, with prompts generated from hand-crafted Radiomics. AutoRad-Lung uses the vision encoder of the Large-Scale Autoregressive Image Model (AIMv2), pre-trained using a multi-modal autoregressive objective. Given that lung tumors are typically small, irregularly shaped, and visually similar to healthy tissue, AutoRad-Lung offers significant advantages over its CLIP-based counterparts by capturing pixel-level differences. Additionally, we introduce conditional context optimization, which dynamically generates context-specific prompts based on input Radiomics, improving cross-modal alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。