用深度学习自动划分放疗部位,解决大数据标注难题。
Automated Dose-Based Anatomic Region Classification of Radiotherapy Treatment for Big Data Applications
- 基于剂量与解剖结构重叠,用深度学习自动推断治疗区域。
- 主治疗区域识别准确率达95%,整体匹配率超91%。
- 适合大规模多中心放疗数据清洗,提升研究可复现性。
放射治疗的大数据应用面临数据清洗瓶颈,尤其当样本量超过10万例时。当前的解剖部位分类依赖不一致的计划标签或靶区命名,难以用于多机构数据。本文开发了一套自动化软件,通过深度学习分割118个解剖结构(器官、腺体、骨骼),结合85%和50%等剂量线生成结构,计算器官特异性剂量重叠指标,据此对治疗计划进行区域标签排序。算法在109例专家标注病例上优化,并在100例临床计划上验证:达到91%精确准确率(完全匹配专家标签及顺序)、94%前两名准确率、95%首项准确率。少数误判多出现在解剖边界区域,属合理模糊范畴。该方法实现高精度、可扩展的标准化数据标注,为放射治疗大数据研究提供可靠支持。
原文摘要 · Abstract (English)
Curation is a significant barrier to using 'big data' radiotherapy planning databases of 100,000+ patients. Anatomic site stratification is essential for downstream analyses, but current methods rely on inconsistent plan labels or target nomenclature, which is unreliable for multi-institutional data. We developed software to automate labeling by inferring anatomic regions directly from dose-volume overlap with deep-learning segmentations, eliminating metadata reliance. The software processes DICOM files in bulk, utilizing deep learning to segment 118 structures (organs, glands, and bones) categorized into six regions: Cranial, Head and Neck, Pelvis, Abdomen, Thorax, Extremity. The 85% and 50% isodose lines are converted to structures to compute organ-specific dose-overlap metrics. Plans are assigned ranked regional labels based on these intersections. The algorithm was refined using 109 expert-labeled cases and validated on 100 consecutive clinical plans. On the 100-plan test dataset, the algorithm achieved 91% Exact Accuracy (matching all expert labels and order), 94% Top-2 Accuracy (matching the top two expert regions regardless of order), and 95% Top-1 Accuracy (matching the primary expert label). The automated workflow demonstrated high accuracy and robustness. The 95% Top-1 Accuracy is particularly significant, as it enables reliable querying of plans based on the primary treatment site. Detailed analysis of the few mismatched cases showed most were treated areas at the border between anatomic regions and were ambiguous between these two regions in a common-sense interpretation. This algorithm provides a scalable, standardized solution for curating the large, multi-institutional datasets required for 'big data' in radiotherapy and provides an important complement to text-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。