arXiv:2605.01266cs.CV2026-05

研究零样本模型如何根据临床信息定位肺癌肿瘤,发现解剖位置最关键。

Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation

论文配图:Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation
图 1 · 摘自论文原文
  • 拆解提示词成分,发现解剖位置决定模型注意力分布
  • 零样本模型平均骰子系数达0.613,接近精细调优模型
  • 适合关注提示词设计与临床信息融合的研究者参考

零样本视觉语言模型(VLMs)为非小细胞肺癌(NSCLC)大体肿瘤体积(GTV)勾画提供了无需任务特定训练的替代方案,但其控制空间行为的提示维度仍不明确。本研究通过在内部NSCLC肿瘤数据集上对VoxTell进行子提示分解(包括诊断、人口统计、分期、解剖、通用及无关控制),属性扰动鲁棒性测试、特异性层级分析和跨病例提示交换实验,评估其性能。采用骰子相似系数(DSC)并结合威尔科克森符号秩检验与邦弗朗尼校正,对比微调及零样本基线。结果显示:解剖位置是主导因素——63.4%的位置扰动导致灾难性性能下降;从通用到完整描述的提示显著提升特异性;无关提示正确生成零分割;跨病例提示交换证实患者特异性条件——匹配时DSC为0.906,不匹配时仅为0.406。组织学与分期替换影响微弱,表明模型更关注‘何处查看’而非‘查看什么’。在此背景下,完全零样本的VoxTell实现均值DSC 0.613,与nnUNet(0.690,调整p=0.156)和Ahmed等(0.675,调整p=0.679)无显著差异,且显著优于其他零样本模型。结果表明,评价分割类VLM应兼顾骰子系数与提示对齐维度。

原文摘要 · Abstract (English)

Zero-shot vision-language models (VLMs) offer a promptable alternative to task-specific training for gross tumor volume (GTV) delineation in non-small-cell lung cancer (NSCLC), but the prompt dimensions that govern their spatial behavior remain poorly understood. We study this question by probing alignment directions in VoxTell on a held-out internal NSCLC tumor dataset through sub-prompt decomposition into diagnosis, demographic, staging, anatomical, generic, and irrelevant controls; attribute-wise perturbation robustness; specificity ladders; and cross-case prompt swaps, while benchmarking against fine-tuned and zero-shot baselines using the Dice Similarity Coefficient (DSC) with Wilcoxon signed-rank tests and Benjamini-Hochberg correction. Alignment analyses revealed that anatomical location is the dominant driver of VoxTell's spatial attention: 63.4 percent of location perturbations caused catastrophic drops, prompt specificity improved from generic to full descriptions except for diagnosis-only prompts, irrelevant prompts correctly yielded zero segmentation, and cross-case prompt swaps confirmed patient-specific conditioning (matched DSC 0.906 vs. mismatched 0.406). Histology and stage substitutions had minimal effect, indicating that the model prioritizes "where to look" over "what to look for." In this context, VoxTell, operating fully zero-shot, achieved a mean DSC of 0.613, statistically indistinguishable from nnUNet (0.690, adjusted p = 0.156) and Ahmed et al. (0.675, adjusted p = 0.679), while significantly outperforming all other zero-shot models. Together, these findings argue that segmentation VLMs should be evaluated not only by Dice, but also by the prompt dimensions to which they align.

零样本分割临床提示解剖定位医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。