构建首个公开胰腺癌血管侵犯评估基准,推动精准术前诊断
Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

- 基于五位专家标注的密集数据集,建立不确定性感知的AI评估框架
- 发现体积重叠高不等于边界判断准,复杂病例中模型易出错
- 适合医学影像AI研究者与临床决策支持系统开发者参考
胰腺导管腺癌(PDAC)手术切除是唯一可能治愈的治疗方式,其可行性依赖于对血管侵犯(VI)的准确评估,即肿瘤是否侵入邻近关键血管。尽管这一评估对术前分期和手术规划至关重要,但计算化VI评估仍缺乏深入研究。主要挑战包括公共数据集匮乏以及肿瘤-血管交界处的诊断模糊性,导致即使在资深放射科医生间也存在显著一致性差异。为此,我们推出了CURVAS-PDACVI数据集与挑战赛,这是一个基于每例扫描五位独立专家标注的开放基准,旨在推动不确定性感知的AI在PDAC分期中的应用。我们还提出了一种多指标评估框架,不仅包含空间重叠,还涵盖概率校准和VI评估。对六种先进方法的评估显示,强全局体积重叠并不保证在临床关键的肿瘤-血管界面表现可靠。特别是,为二值分割优化的方法在平均重叠指标上表现良好,但在低共识的高复杂度病例中常出现体积坍缩或边界过度延伸。相比之下,建模医师范围分歧的方法生成了更校准的概率图,在这些模糊案例中表现出更强鲁棒性。该基准揭示了体积精度作为局部手术实用性代理的局限性,推动采用不确定性感知的概率模型以支持术前决策。
原文摘要 · Abstract (English)
Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate assessment of vascular invasion (VI), i.e., tumor extension into adjacent critical vessels. Despite its importance for preoperative staging and surgical planning, computational VI assessment remains underexplored. Two major challenges are the lack of public datasets and the diagnostic ambiguity at the tumor-vessel interface, which leads to substantial inter-rater variability even among expert radiologists. To address these limitations, we introduce the CURVAS-PDACVI Dataset and Challenge, an open benchmark for uncertainty-aware AI in PDAC staging based on a densely annotated dataset with five independent expert annotations per scan. We also propose a multi-metric evaluation framework that extends beyond spatial overlap to include probabilistic calibration and VI assessment. Evaluation of six state-of-the-art methods shows that strong global volumetric overlap does not necessarily translate into reliable performance at clinically critical tumor-vessel interfaces. In particular, methods optimized for binary segmentation perform competitively on average overlap metrics, but often degrade in high-complexity cases with low expert consensus, either collapsing in volume or overextending at uncertain boundaries. In contrast, methods that model inter-rater disagreement produce better calibrated probabilistic maps and show greater robustness in these ambiguous cases. The benchmark highlights the limitations of volumetric accuracy as a proxy for localized surgical utility, motivating uncertainty-aware probabilistic models for preoperative decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。