arXiv:2608.19813eess.IV2026-08中稿 · the MICCAI 2026 wo…

用物理约束提升医学图像分类的效率与精度

Energy-Mamba: A Physics-Constrained State-Space Model for Medical Image Classification

论文配图:Energy-Mamba: A Physics-Constrained State-Space Model for Medical Image Classification
图 1 · 摘自论文原文
  • 引入可学习势能函数,约束状态演化保持局部特征清晰
  • 在4个医学图像数据集上达顶尖性能,参数量显著减少
  • 适合对细节敏感的医学视觉任务,如病理诊断

状态空间模型(SSMs)尤其是Mamba,因其线性时间复杂度在长序列建模中表现优异,适用于标注数据有限的医学影像任务。然而,通过无约束状态演化处理2D图像时,隐藏状态会逐渐偏离原始局部特征,导致表征漂移。本文提出Energy-Mamba,将SSM动态与物理启发的约束结合,引入可学习的势能函数,量化状态演化与静态局部特征的兼容性。能量块通过自动微分计算梯度驱动项,动态引导状态向低能量配置靠拢,以维持局部视觉保真度。该设计模仿哈密顿动力学:动能(SSM扫描)+ 势能(约束函数)共同决定状态轨迹。这种架构先验使模型能隐式学习稳健的约束,提升表示质量,在四组数据集(视网膜OCT、胸部X光、显微镜、腹部CT)上均达到当前最优分类性能,且参数量大幅减少,验证了物理信息引导在医学视觉任务中同时提升效率与表征质量的有效性。

原文摘要 · Abstract (English)

State-Space Models (SSMs), particularly Mamba, offer linear-time complexity for long-range dependencies, making them attractive for medical imaging with limited annotated data. However, adapting these sequential models to 2D images through unconstrained state evolution causes representational drift, the dynamic hidden state progressively loses fidelity to local image features. We introduce Energy-Mamba, integrating SSM dynamics with physics-informed constraints via a learnable potential energy function that quantifies compatibility between evolving states and static local features. Our Energy-Mamba Block introduces a gradient-based forcing term, computed dynamically via automatic differentiation, that pulls states toward low-energy configurations maintaining local visual fidelity. This formulation mirrors Hamiltonian dynamics: kinetic energy (SSM scan) plus potential energy (our constraint function) govern state trajectories. This architectural prior enables learning implicit constraints for robust, faithful representations, crucial in medical imaging where fine-grained local detail drives accurate diagnosis. Evaluated on four datasets (retinal OCT, chest X-ray, microscopy, abdominal CT), Energy-Mamba achieves state-of-the-art classification performance with significantly fewer parameters, demonstrating that physics-informed grounding can enhance both efficiency and representational quality in medical vision tasks.

医学图像状态空间模型物理约束高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。