arXiv:2604.10702cs.CVcs.AI2026-04被引 1

不同主干网络下,模态门控行为差异显著,Mamba更稳定。

Backbone-Conditional Behavior of Modality Gating in Multi-Modal Prostate MRI Segmentation: A 5-Fold Cross-Validation and Gate Mechanism Analysis

论文配图:Backbone-Conditional Behavior of Modality Gating in Multi-Modal Prostate MRI Segmentation: A 5-Fold Cross-Validation and Gate Mechanism Analysis
图 1 · 摘自论文原文
  • 对比nnU-Net与Mamba,发现门控行为受主干网络影响
  • Mamba的门控保持样本级变化,提升跨队列鲁棒性
  • 训练时模态丢弃是唯一在两类主干上均有效的策略

多参数MRI中对临床显著前列腺癌(csPCa)的鲁棒分割需应对最信息量的扩散序列频繁退化问题。现有方法依赖学习式模态门控,假设其能实现按样本的模态质量路由,但极少被直接验证。本文系统分析了在nnU-Net与Mamba两个主干上,基于PI-CAI(n=1500)的模态隔离门控融合(MIGF),并在Prostate158(n=158)进行跨队列验证。通过5折交叉验证共训练180个模型,结合门控、模态丢弃与深度监督的因子设计,并对30个门控模型进行门权重与反事实分析。结果显示:门控行为具有主干依赖性——在nnU-Net上,加入门控使排名分数下降(边际效应-0.037,p<0.05);而在Mamba上,门控+丢弃配置则提升分数(+0.024,p=0.037)。门权重分析表明:nnU-Net的门控趋近静态模态先验(跨案例标准差0.0033),而Mamba保留样本级动态(0.0365,约11倍大,无重叠);用训练集均值替代样本门控,nnU-Net不变,但Mamba性能下降。模态丢弃是唯一在两主干上均有益的组件。跨队列迁移中,卷积主干性能降至近零特异性,而Mamba维持较高水平(最高达0.31)。结论:学习式门控并非普遍实现样本级质量路由;其有效行为取决于主干自身的模态感知能力。在测试配置中,MIGF-Mamba最具跨队列鲁棒性,且训练时模态丢弃是唯一通用增益项。

原文摘要 · Abstract (English)

Robust segmentation of clinically significant prostate cancer (csPCa) on multi-parametric MRI must tolerate frequent degradation of its most informative diffusion sequences. Multi-modal fusion commonly employs learned modality gating under the assumption that gates implement per-sample modality quality routing -- rarely tested directly. We ask how gating behaves across backbone architectures. We systematically analyze modality-isolated gated fusion (MIGF) for csPCa segmentation on two backbones (nnU-Net and Mamba) using PI-CAI (n=1500), with cross-cohort validation on Prostate158 (n=158): a factorial ablation over gating, modality dropout, and deep supervision under 5-fold cross-validation (180 trained models), plus a gate-weight and counterfactual analysis of 30 trained gating models. Modality gating is backbone-conditional. On nnU-Net, adding gating reduces the ranking score (marginal effect -0.037; gating configurations p<0.05), whereas on Mamba the gating-plus-dropout configuration improves it (+0.024, p=0.037). Gate-weight analysis explains this: nnU-Net gates collapse into a near-static modality prior (across-case SD 0.0033), while Mamba gates retain sample-dependent variation (0.0365, ~11x larger, non-overlapping); replacing per-sample gates with their training-set mean leaves nnU-Net unchanged but degrades Mamba. Modality dropout is the only component beneficial on both backbones. Under cross-cohort shift, convolutional backbones collapse to case-level specificity near zero, whereas Mamba retains it (MIGF-Mamba highest, 0.31). Learned modality gates do not universally perform per-sample quality routing; their effective behavior is conditional on the backbone's inherent modality awareness. Among tested configurations, MIGF-Mamba is the most cross-cohort robust, and training-time modality dropout is the only component beneficial across both backbones.

医学图像多模态门控机制鲁棒分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。