arXiv:2601.09130eess.IVcs.AI2026-01中稿 · IEEE ISBI 2026 4-p…

让视觉Transformer具备旋转不变性,提升病理图像分析稳定性

Equi-ViT: Rotational Equivariant Vision Transformer for Robust Histopathology Analysis

  • 在图像嵌入阶段引入等变卷积,使模型对旋转保持一致
  • 在结直肠癌数据集上实现更稳定分类性能,提升数据利用效率
  • 适合构建鲁棒的数字病理学基础模型,尤其关注旋转变化场景

视觉变换器(ViT)因其自注意力机制可捕捉长距离依赖,在计算病理学中迅速普及,弥补了传统卷积网络仅擅长局部模式而难以处理全局上下文的缺陷。近期病理专用基础模型通过大规模预训练进一步提升了性能。然而,标准ViT对旋转、翻转等变换不具备等变性,而这些变换在病理图像中普遍存在。为此,我们提出Equi-ViT,将等变卷积核集成至ViT的补丁嵌入阶段,赋予学习表征内在的旋转等变性。Equi-ViT实现了更优的旋转一致性补丁嵌入与跨图像方向的稳定分类表现。在公开结直肠癌数据集上的实验表明,引入等变嵌入显著提升数据效率与鲁棒性,提示等变变压器或可作为更通用的骨干网络应用于病理学中的ViT,如数字病理基础模型。

原文摘要 · Abstract (English)

Vision Transformers (ViTs) have gained rapid adoption in computational pathology for their ability to model long-range dependencies through self-attention, addressing the limitations of convolutional neural networks that excel at local pattern capture but struggle with global contextual reasoning. Recent pathology-specific foundation models have further advanced performance by leveraging large-scale pretraining. However, standard ViTs remain inherently non-equivariant to transformations such as rotations and reflections, which are ubiquitous variations in histopathology imaging. To address this limitation, we propose Equi-ViT, which integrates an equivariant convolution kernel into the patch embedding stage of a ViT architecture, imparting built-in rotational equivariance to learned representations. Equi-ViT achieves superior rotation-consistent patch embeddings and stable classification performance across image orientations. Our results on a public colorectal cancer dataset demonstrate that incorporating equivariant patch embedding enhances data efficiency and robustness, suggesting that equivariant transformers could potentially serve as more generalizable backbones for the application of ViT in histopathology, such as digital pathology foundation models.

视觉Transformer病理图像等变性数字病理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。