arXiv:2511.19367cs.CVcs.AI2025-11被引 1

将肺癌分期转化为可解释的测量与规则推理,提升临床可信度。

AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging

  • 分三步:分割肺、肿瘤、纵隔,用轮廓法测尺寸与距离
  • 准确率91.36%,各期F1值达0.89~0.96,优于以往方法
  • 适合需要可解释性决策的临床深度学习场景

肺癌精准分期对预后与治疗至关重要,其依据是明确的解剖学标准。然而现有深度学习方法多将其视为不可解释的图像分类任务。肿瘤分期依赖于定量标准,包括肿瘤大小及其与邻近解剖结构的距离,微小差异即可改变分期结果。为此,我们提出AnatomicalNets,一个医学基准的多阶段流程,将分期重构为测量与规则推理问题。采用三个专用编码器-解码器网络精确分割肺实质、肿瘤和纵隔;通过肺轮廓启发式估计膈肌边界;利用基于轮廓的距离估计法计算肿瘤最大直径及其与邻近结构的距离。这些特征输入遵循国际肺癌研究协会指南的确定性决策模块。在Lung-PET-CT-Dx数据集上,AnatomicalNets整体分类准确率达91.36%。各期F1分数分别为:T1为0.93,T2为0.89,T3为0.96,T4为0.90,这一关键评估指标此前常被忽略。我们指出,先前工作的表征瓶颈在于特征设计而非分类器容量。本工作建立了一种透明可靠的分期范式,弥合了深度学习性能与临床可解释性之间的差距。

原文摘要 · Abstract (English)

Accurate tumor staging in lung cancer is crucial for prognosis and treatment planning and is governed by explicit anatomical criteria under fixed guidelines. However, most existing deep learning approaches treat this spatially structured clinical decision as an uninterpretable image classification problem. Tumor stage depends on predetermined quantitative criteria, including the tumor's dimensions and its proximity to adjacent anatomical structures, and small variations can alter the staging outcome. To address this gap, we propose AnatomicalNets, a medically grounded, multi-stage pipeline that reformulates tumor staging as a measurement and rule-based inference problem rather than a learned mapping. We employ three dedicated encoder-decoder networks to precisely segment the lung parenchyma, tumor, and mediastinum. The diaphragm boundary is estimated via a lung-contour heuristic, while the tumor's largest dimension and its proximity to adjacent structures are computed through a contour-based distance estimation method. These features are passed through a deterministic decision module following the international association for the study of lung cancer guidelines. Evaluated on the Lung-PET-CT-Dx dataset, AnatomicalNets achieves an overall classification accuracy of 91.36%. We report the per-stage F1-scores of 0.93 (T1), 0.89 (T2), 0.96 (T3), and 0.90 (T4), a critical evaluation aspect often omitted in prior literature. We highlight that the representational bottleneck in prior work lies in feature design rather than classifier capacity. This work establishes a transparent and reliable staging paradigm that bridges the gap between deep learning performance and clinical interpretability.

医学影像分割可解释性肺癌分期

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。