arXiv:2605.05522eess.IVcs.CV2026-05中稿 · Biomedical Signal …

针对跨模态医学图像分割中的模型失效问题,提出两种针对性改进策略。

Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images

论文配图:Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images
图 1 · 摘自论文原文
  • 通过肿瘤感知增强和非对称裁剪提升模型对MRI数据的适配效率。
  • 在直肠癌MRI分割任务中,最高检测率达91.1%(225/247)。
  • 揭示了预训练模型跨模态迁移的两大失败机制,适合医学影像研究者参考。

尽管自监督预训练有望学习广泛可迁移的表征,但其在与预训练域差异较大的成像模态上,以及复杂肿瘤分割任务中的有效性仍缺乏研究。评估基于CT预训练的Transformer在MRI直肠癌分割中的表现时,我们识别出两个相互作用的失败模式:(a) 因零填充以匹配预训练输入尺寸导致的令牌使用效率低下;(b) 特征适应无效。我们利用两种主要的基于CT预训练的分层移位窗口变压器主干网络SMIT和Swin UNETR,以及VoCo作为大规模预训练基准进行研究;这些模型在预训练目标和数据集上存在差异。机制分析采用注意力稀释指数(ADI),一种基于熵的度量,用于量化注意力向无信息填充令牌的偏移;并使用中心核对齐(CKA)衡量在MRI适应过程中的特征复用情况。结果显示,零填充导致ADI上升,而高特征复用并不必然带来下游准确率提升。为缓解这些问题,我们引入两项干预措施:肿瘤感知增强策略以扩大肿瘤外观异质性的覆盖范围,以及各向异性裁剪策略以恢复令牌效率。在相同直肠MRI数据集上微调后,主干网络SMIT和Swin UNETR的检测率分别达到91.1%(225/247)和88.7%(219/247),支持性基准VoCo达到90.3%(223/247),表明在跨模态转移下显著提升鲁棒性。本研究是首批系统探讨预训练变换器在跨模态迁移中失效原因的研究之一,并展示了针对性缓解策略如何有效克服跨模态迁移限制。

原文摘要 · Abstract (English)

Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities substantially different from the pretraining domain, and on complex tumor-segmentation tasks, remains understudied. Evaluating CT-pretrained transformers on MRI rectal cancer segmentation, we identified two interacting failure modes in CT-to-MRI transfer: (a) inefficient token usage caused by zero-padding to match pretrained input dimensions, and (b) ineffective feature adaptation. We investigated these vulnerabilities using two primary CT-pretrained hierarchical shifted-window transformer backbones, SMIT and Swin UNETR, together with VoCo as a large-scale-pretrained supporting benchmark; these models differ in pretraining objectives and datasets. Mechanistic analysis leveraged an attention dilution index (ADI), an entropy-based metric quantifying attention diverted toward uninformative padding tokens, and centered kernel alignment (CKA) to measure feature reuse during MRI adaptation. ADI increased with zero-padding, while high feature reuse did not necessarily translate to improved downstream accuracy. To mitigate these issues, we introduced two interventions: a tumor-aware augmentation strategy to expand tumor appearance heterogeneity coverage, and an anisotropic cropping strategy to restore token efficiency. Fine-tuning with these strategies on identical rectal MRI datasets yielded detection rates of 91.1% (225/247) and 88.7% (219/247) for the primary SMIT and Swin UNETR backbones, with the supporting VoCo benchmark reaching 90.3% (223/247), demonstrating significantly improved robustness under CT-to-MRI transfer. This study is among the first to examine when pretrained transformers fail to transfer across imaging modalities and demonstrates how targeted mitigation strategies can systematically overcome cross-modality transfer limitations.

医学图像跨模态分割Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。