arXiv:2512.03883cs.CV2025-12中稿 · ISBI 2026 conferen…

用双注意力网络分析肠镜图像,早判肿瘤是否复发。

Dual Cross-Attention Siamese Transformer for Rectal Tumor Regrowth Assessment in Watch-and-Wait Endoscopy

  • 设计双交叉注意力结构,无需对齐图像即可比对前后扫描
  • 在62例患者数据上实现90%灵敏度、81.76%准确率
  • 对血迹、粪便等干扰有强鲁棒性,适合临床实用

越来越多证据支持对经新辅助治疗后达到临床完全缓解(cCR)的直肠癌患者实施观察等待(WW)策略。然而,在随访肠镜中早期准确检测局部复发(LR)对管理治疗和预防远处转移至关重要。为此,我们提出一种基于双交叉注意力的孪生Swin Transformer(SSDCA),通过结合复诊与随访的肠镜图像,区分cCR与LR。SSDCA利用预训练Swin Transformer提取与领域无关的特征,增强对成像差异的鲁棒性。双交叉注意力机制强调配对扫描中的关键特征,无需任何空间对齐即可预测疗效。模型在135名患者的图像对上训练,于62名患者独立测试集上评估:平衡准确率81.76%±0.04,敏感度90.07%±0.08,特异度72.86%±0.05。鲁棒性分析显示,即使存在血迹、粪便、毛细血管扩张或图像质量差,性能仍稳定。特征的UMAP聚类显示,SSDCA实现最大簇间分离(1.45±0.18)与最小簇内分散(1.07±0.19),验证其具备强判别性表示能力。代码与权重已公开:https://github.com/Jotanator/SSDCA

原文摘要 · Abstract (English)

Increasing evidence supports watch-and-wait (WW) surveillance for patients with rectal cancer who show clinical complete response (cCR) at restaging following total neoadjuvant treatment (TNT). However, accurate methods to early detect local regrowth (LR) from follow-up endoscopy images during WW are essential to manage care and prevent distant metastases. Hence, we developed a Siamese Swin Transformer with Dual Cross-Attention (SSDCA) to combine longitudinal endoscopic images at restaging and follow-up and distinguish cCR from LR. SSDCA leverages pretrained Swin Transformers to extract domain agnostic features and enhance robustness to imaging variations. Dual cross attention is implemented to emphasize features from the paired scans without requiring any spatial alignment to predict response. SSDCA as well as Swin-based baselines were trained using image pairs from 135 patients and evaluated on a held-out set of image pairs from 62 patients. SSDCA produced the best balanced accuracy (81.76% $\pm$ 0.04), sensitivity (90.07% $\pm$ 0.08), and specificity (72.86% $\pm$ 0.05). Robustness analysis showed stable performance irrespective of artifacts including blood, stool, telangiectasia, and poor image quality. UMAP clustering of extracted features showed maximal inter-cluster separation (1.45 $\pm$ 0.18) and minimal intra-cluster dispersion (1.07 $\pm$ 0.19) with SSDCA, confirming discriminative representation learning. Code and weights available at: https://github.com/Jotanator/SSDCA

医学影像图像对比深度学习肠镜分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。