arXiv:2412.13294cs.CVcs.LG2024-12

提出可解释的可变形图像配准方法,提升脑部与视网膜图像配准精度与可理解性。

Interpretable deformable image registration: A geometric deep learning perspective

  • 分离特征提取与形变建模,利用几何深度学习实现连续空间形变。
  • 在多模态脑图像与纵向视网膜图像上,性能超越现有最佳方法。
  • 支持从粗到精的端到端优化,避免网格重采样带来的误差。

可变形图像配准面临复杂坐标系统间关系建模的挑战。尽管数据驱动方法展现出强大非线性变换建模能力,但现有工作多采用标准深度学习架构,将其视为通用黑箱求解器。本文认为,理解模型如何在源域与目标域特征间进行模式匹配,是构建鲁棒、数据高效且可解释架构的关键。我们提出了一个理论基础:分离的特征提取与形变建模、动态感受野,以及对双空间关系的数据驱动感知。基于此,设计了一种端到端的粗到精优化流程。架构采用基于几何深度学习原理的连续空间形变函数,避免了在迭代过程中对齐至规则网格的繁琐重采样操作。通过定性分析揭示了模型良好的可解释性特性。结果表明,在单模态与多模态跨被试脑图像配准,以及纵向视网膜内被试配准等挑战性任务中,本方法显著优于当前最先进水平。代码已公开。

原文摘要 · Abstract (English)

Deformable image registration poses a challenging problem where, unlike most deep learning tasks, a complex relationship between multiple coordinate systems has to be considered. Although data-driven methods have shown promising capabilities to model complex non-linear transformations, existing works employ standard deep learning architectures assuming they are general black-box solvers. We argue that understanding how learned operations perform pattern-matching between the features in the source and target domains is the key to building robust, data-efficient, and interpretable architectures. We present a theoretical foundation for designing an interpretable registration framework: separated feature extraction and deformation modeling, dynamic receptive fields, and a data-driven deformation functions awareness of the relationship between both spatial domains. Based on this foundation, we formulate an end-to-end process that refines transformations in a coarse-to-fine fashion. Our architecture employs spatially continuous deformation modeling functions that use geometric deep-learning principles, therefore avoiding the problematic approach of resampling to a regular grid between successive refinements of the transformation. We perform a qualitative investigation to highlight interesting interpretability properties of our architecture. We conclude by showing significant improvement in performance metrics over state-of-the-art approaches for both mono- and multi-modal inter-subject brain registration, as well as the challenging task of longitudinal retinal intra-subject registration. We make our code publicly available

图像配准几何深度学习可解释性脑影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。