用强化学习实现快速精准的术前CT与术中腹腔镜图像配准。
Warm-Started Reinforcement Learning for Iterative 3D/2D Liver Registration

- 将图像配准建模为连续决策过程,通过强化学习自动选择变换并决定停止时机。
- 在公开数据集上平均配准误差达15.70毫米,媲美传统优化方法且收敛更快。
- 基于预训练网络初始化,提升稳定性与效率,适合手术增强现实场景。
术前CT与术中腹腔镜视频之间的配准在微创手术的增强现实导航中至关重要。基于学习的方法近年来实现了与优化方法相当的配准误差,同时具备更快的推理速度。然而,许多监督方法生成粗略对齐,需依赖额外优化步骤进行精化,从而增加推理时间。本文提出一种离散动作强化学习(RL)框架,将CT到视频的配准建模为序列决策过程。共享特征编码器从预训练的监督姿态估计网络中获取初始化,提供稳定几何特征并加速收敛,分别从CT渲染图和腹腔镜帧中提取表示;而强化学习策略头则学习在六自由度上选择刚性变换,并决定何时停止迭代。在公开腹腔镜数据集上的实验表明,该方法平均目标配准误差(TRE)为15.70毫米,与结合优化的监督方法相当,同时实现更快收敛。所提出的基于强化学习的框架实现了无需人工调参步长或停止条件的自动化、高效迭代配准,为未来外科增强现实中连续动作与可变形配准模型提供了实用基础。
原文摘要 · Abstract (English)
Registration between preoperative CT and intraoperative laparoscopic video plays a crucial role in augmented reality (AR) guidance for minimally invasive surgery. Learning-based methods have recently achieved registration errors comparable to optimization-based approaches while offering faster inference. However, many supervised methods produce coarse alignments that rely on additional optimization-based refinement, thereby increasing inference time. We present a discrete-action reinforcement learning (RL) framework that formulates CT-to-video registration as a sequential decision-making process. A shared feature encoder, warm-started from a supervised pose estimation network to provide stable geometric features and faster convergence, extracts representations from CT renderings and laparoscopic frames, while an RL policy head learns to choose rigid transformations along six degrees of freedom and to decide when to stop the iteration. Experiments on a public laparoscopic dataset demonstrated that our method achieved an average target registration error (TRE) of 15.70 mm, comparable to supervised approaches with optimization, while achieving faster convergence. The proposed RL-based formulation enables automated, efficient iterative registration without manually tuned step sizes or stopping criteria. This discrete framework provides a practical foundation for future continuous-action and deformable registration models in surgical AR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。