arXiv:2608.11076cs.CV2026-08被引 1

用大模型选最值得标注的癌症影像,少标30%数据仍达全量标注效果。

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

论文配图:Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets
图 1 · 摘自论文原文
  • 基于视觉大模型嵌入,一键挑选最具代表性的未标注病例
  • 仅需70%标注量即达全量标注模型性能,跨癌种跨示踪剂通用
  • 适合临床科研中数据标注成本高的场景,尤其适用于多模态医学影像

全身影像组学中自动病灶分割可辅助临床进行癌症检测、分期与治疗规划,但其模型训练需大量标注数据,而标注耗时耗力。本文提出FEEDS(Foundation model-Enabled Efficient Data Sampling)策略,利用视觉基础模型嵌入,从海量未标注数据中高效筛选最具信息量和多样性病例供专家标注。该方法为单次训练范式,仅需少量代表性数据即可,显著降低标签与算力需求。在AutoPET-III数据集上训练并验证,测试于三个独立数据集:AutoPET-III、DeepPSMA及达特茅斯-希奇科克医疗中心内部数据集。在体素、病灶及解剖区域层面评估临床实用性,结果表明其性能优于随机采样、伪标签半监督学习及仅用有限标注数据训练的方法,且在三种测试集、FDG与PSMA示踪剂、多种疾病间均表现良好,仅需70%标注量即可达到100%全标注训练的性能水平。该方法有效缓解了自动病灶分割中的标签稀缺问题,为构建高质量标注队列提供可行路径。

原文摘要 · Abstract (English)

Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires large annotated datasets, whose creation is both time- and expertise-intensive. As a result, models trained on limited labeled PET/CT data often lack the accuracy and generalizability needed for clinical use. We present FEEDS (Foundation model-Enabled Efficient Data Sampling), a label- and compute-efficient learning strategy that uses vision foundation model embeddings to select the most informative and diverse unlabeled cases for expert annotation. Unlike unsupervised, semi-supervised, and active learning approaches, FEEDS is a one-step training paradigm requiring only a limited, representative training set, making it label- and compute-efficient. We train and validate FEEDS using the AutoPET-III dataset. We test its accuracy and generalizability on three held-out sets: AutoPET-III, DeepPSMA, and an internal Dartmouth-Hitchcock Medical Center dataset. We evaluate clinical utility at the voxel, lesion, and anatomic region level to assess performance in high-risk areas and treatment planning utility. FEEDS outperforms random-sampling-based labeling, pseudolabel-based semi-supervised learning, and training with limited labeled data alone. It generalizes across all three test sets, FDG and PSMA tracers, and multiple diseases, matching fully-labeled (100\%) training performance with 70\% less annotation burden. FEEDS addresses the challenge of label scarcity in an automatic lesion segmentation framework by providing a practical approach for constructing representative and diverse annotation queues from large, unannotated clinical repositories.

医学影像小样本学习大模型应用癌症诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。