用预训练Transformer实现少样本心脏结构分割,提升放疗规划精度与泛化能力。
Transformer-based cardiac substructure segmentation from contrast and non-contrast computed tomography for radiotherapy planning
- 基于混合预训练Transformer-卷积网络,结合均衡课程学习策略。
- 仅用64%数据达到与全量数据相当的分割精度(HD95: 6.6 vs 5.4mm)。
- 对不同患者体位和扫描协议均表现稳健,适合临床多场景应用。
准确分割CT影像中的心脏亚结构对放疗计划至关重要,但通常需大量标注数据且跨成像协议和患者差异时泛化能力差。本研究评估了预训练Transformer在固定架构下实现数据高效训练的可行性,采用混合预训练Transformer-卷积网络SMIT,在180例肺癌患者(仰卧位)上微调,并在60例保留的同类患者及65例乳腺癌患者(仰卧/俯卧位)上验证。对比两种配置:SMIT-Balanced(32例增强+32例非增强CT)与SMIT-Oracle(180例CT)。性能以95%分位数豪斯多夫距离(HD95)为主指标,辐射剂量与重叠度为辅。SMIT-Balanced仅用64%训练数据即达与SMIT-Oracle相当精度(同组内:6.6±4.3 mm vs 5.4±2.6 mm;跨组:10.0±9.4 mm vs 9.4±9.8 mm),对患者、成像及数据变化均具鲁棒性。基于SMIT分割的剂量计算结果与人工勾画一致。尽管nnU-Net优于公开模型TotalSegmentator,但其跨域泛化能力弱于SMIT。均衡课程学习在降低标注需求的同时保持精度,并减少对特定数据域的依赖,避免了nnU-Net所需的数据定制架构调整。
原文摘要 · Abstract (English)
Accurate segmentation of cardiac substructures on computed tomography (CT) scans is essential for radiotherapy planning but typically requires large annotated datasets and often generalizes poorly across imaging protocols and patient variations. This study evaluated whether pretrained transformers enable data-efficient training using a fixed architecture with balanced curriculum learning. A hybrid pretrained transformer-convolutional network (SMIT) was fine-tuned on lung cancer patients (Cohort I, N $=$ 180) imaged in the supine position and validated on 60 held-out Cohort I patients and 65 breast cancer patients (Cohort II) imaged in both supine and prone positions. Two configurations were evaluated: SMIT-Balanced (32 contrast-enhanced CTs and 32 non-contrast CTs) and SMIT-Oracle (180 CTs). Performance was compared with nnU-Net and TotalSegmentator. Segmentation accuracy was assessed primarily using the 95th percentile Hausdorff distance (HD95), with radiation dose and overlap-based metrics evaluated as secondary endpoints. SMIT-Balanced achieved accuracy comparable to SMIT-Oracle despite using 64$\%$ fewer training scans. On Cohort I, HD95 was 6.6 $\pm$ 4.3 mm versus 5.4 $\pm$ 2.6 mm, and on Cohort II, 10.0 $\pm$ 9.4 mm versus 9.4 $\pm$ 9.8 mm, respectively, demonstrating robustness to patient, imaging, and data variations. Radiation dose metrics derived from SMIT segmentations were equivalent to those from manual delineations. Although nnU-Net improved over the publicly trained TotalSegmentator, it showed reduced cross-domain robustness compared to SMIT. Balanced curriculum training reduced labeled data requirements without compromising accuracy relative to the oracle model and maintained robustness across patient and imaging variations. Pretraining reduced dependence on data domain and obviated the need for data-specific architectural reconfiguration required by nnU-Net.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。