arXiv:2511.17392cs.CV2025-11

用细粒度潜空间策略优化,提升医学图像形变配准精度与效率

MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration

  • 在潜空间构建随机高斯策略,实现从粗到精的连续优化
  • 三组3D医学影像数据上Dice评分均优于现有方法
  • 无需密集标注,参数量小,适合高维视觉对齐任务

可变形图像配准(DIR)在医学图像分析中至关重要但挑战重重,主要源于密集位移场的高维形变空间和稀疏体素级监督。现有强化学习框架常将该空间投影为粗糙低维表示,难以捕捉空间异质性形变。本文提出MorphSeek,一种细粒度潜空间策略优化范式,将DIR重构为潜特征空间中的空间连续优化过程。通过在编码器上引入随机高斯策略头,建模潜特征分布,实现高效探索与粗到精迭代优化。框架结合无监督预训练与弱监督微调,采用分组相对策略优化(Group Relative Policy Optimization),多轨迹采样稳定训练并提升标签效率。在三个3D配准基准(OASIS脑部MRI、LiTS肝脏CT、Abdomen MR-CT)上,MorphSeek持续优于竞争基线,保持高标签效率,参数开销极小且每步延迟低。MorphSeek不仅突破优化器局限,更推动了表示级策略学习范式的发展,实现空间一致、数据高效的形变优化,为高维场景下的可扩展视觉对齐提供了一种原则性强、模型无关、优化器无关的解决方案。

原文摘要 · Abstract (English)

Deformable image registration (DIR) remains a fundamental yet challenging problem in medical image analysis, largely due to the prohibitively high-dimensional deformation space of dense displacement fields and the scarcity of voxel-level supervision. Existing reinforcement learning frameworks often project this space into coarse, low-dimensional representations, limiting their ability to capture spatially variant deformations. We propose MorphSeek, a fine-grained representation-level policy optimization paradigm that reformulates DIR as a spatially continuous optimization process in the latent feature space. MorphSeek introduces a stochastic Gaussian policy head atop the encoder to model a distribution over latent features, facilitating efficient exploration and coarse-to-fine refinement. The framework integrates unsupervised warm-up with weakly supervised fine-tuning through Group Relative Policy Optimization, where multi-trajectory sampling stabilizes training and improves label efficiency. Across three 3D registration benchmarks (OASIS brain MRI, LiTS liver CT, and Abdomen MR-CT), MorphSeek achieves consistent Dice improvements over competitive baselines while maintaining high label efficiency with minimal parameter cost and low step-level latency overhead. Beyond optimizer specifics, MorphSeek advances a representation-level policy learning paradigm that achieves spatially coherent and data-efficient deformation optimization, offering a principled, backbone-agnostic, and optimizer-agnostic solution for scalable visual alignment in high-dimensional settings.

图像配准强化学习潜空间医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。