多维注意力提升胃肠道内镜分类,但效果依赖数据分布差异。
MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

- 设计并行的通道、空间和上下文注意力头,适配不同数据分布。
- 在大差异数据集上提升6个模型,最高宏F1达98.33%,小差异数据集则无效。
- 注意力组合有效,单独使用未必有益,适合跨域胃肠道图像分类任务。
自动化胃肠道(GI)内镜分类需模型在多样模态和类别分布间具备泛化能力,常偏离自然图像预训练分布。我们提出MultiAttenGastro,一种即插即用的多维注意力框架,包含并行的一维通道、二维空间与三维上下文注意力头,并首次系统性评估了八种CNN与Transformer骨干网络在五个公开胃肠道数据集上的表现(共80次骨干-数据集组合)。结果表明,注意力有效性并非普适,而是与ImageNet特征与目标分布间的表征差距相关:MultiAttenGastro在代表大差距的Kvasir-Capsule(14类胶囊内镜,宏F1最高98.33%)上改善6个骨干网络,但在小差距的Kvasir-v2基准上全为负向(0/8),中间差距数据集表现混合。五次种子消融实验显示,该改进方向一致但统计不显著(配对t检验:p=0.47;Wilcoxon:p=0.63),且单个注意力头孤立使用并无普遍增益,仅组合后呈现正向均值效应。中心核对齐(CKA)分析揭示此模式与表征冗余有关:大领域差距下头间低互相关性对应一致收益,小差距下高冗余性对应性能下降。我们报告这些结果,包括非显著提升,作为多维注意力在何时何地有效的实证依据,而非宣称其为绝对更优架构。
原文摘要 · Abstract (English)
Automated gastrointestinal (GI) endoscopy classification requires models that generalize across diverse modalities and class distributions, often far from natural-image pretraining. We propose MultiAttenGastro, a plug-and-play attention framework with parallel 1-D channel, 2-D spatial, and 3-D contextual heads, and present the first systematic cross-dataset evaluation across eight CNN and transformer backbones on five public GI datasets (80 backbone--dataset runs). We find that attention effectiveness is not universal but tracks the representational gap between ImageNet features and the target distribution: MultiAttenGastro improves 6 of 8 backbones on Kvasir-Capsule (14-class WCE, large gap; best macro F1 98.33\%), is uniformly negative on the small-gap Kvasir-v2 benchmark (0/8), and shows mixed outcomes on datasets with intermediate gap. Five-seed ablation on the strongest case (Kvasir-Capsule, ConvNeXt-Tiny) shows this improvement is directionally consistent, but not statistically decisive (paired $t$: $p=0.47$; Wilcoxon: $p=0.63$), and that individual attention heads are not uniformly beneficial in isolation only their combination yields a positive mean effect. Centered Kernel Alignment (CKA) analysis links this pattern to representational redundancy: low inter-head CKA under large domain gaps coincides with the framework's only consistent gains, while high redundancy under small gaps coincides with its losses. We report these results, including the non-significant margins, as evidence for when and why multi-dimensional attention helps GI endoscopy classification, rather than as a claim that MultiAttenGastro is a strictly superior architectural choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。