用学习的运动策略预测内镜相机位姿,更稳定且适应低纹理环境。
Investigating a Policy-Based Formulation for Endoscopic Camera Pose Recovery
- 基于专家经验设计运动策略,直接预测短时相对位移。
- 在尸体鼻窦手术数据上实现最低平移误差和良好旋转精度。
- 对低纹理和光照变化更鲁棒,适合临床实时导航场景。
在内镜手术中,外科医生通过结合先前知识与术中视觉动态,持续定位内镜视角与解剖结构的关系。现有视觉导航系统试图复现这一能力,但大多依赖关键帧间的特征匹配与几何优化,难以应对内镜成像中的低纹理、快速光照变化等挑战。本文提出一种基于策略的内镜相机位姿恢复方法,模仿专家根据前一时刻状态推断轨迹的推理方式。该方法不维护显式几何表示,直接预测短时相对运动,从设计上缓解了传统几何方法中对应关系脆弱、纹理稀疏区域不稳定及重建失败导致位姿覆盖有限等问题。在尸体鼻窦内镜数据上评估,采用理想状态条件时,其短时运动预测在平移误差上优于几何基线,旋转精度具有竞争力;按纹理丰富度与光照变化分组分析显示,对低纹理条件敏感性更低。结果表明,学习的运动策略是内镜位姿恢复的一种可行替代方案。
原文摘要 · Abstract (English)
In endoscopic surgery, surgeons continuously locate the endoscopic view relative to the anatomy by interpreting the evolving visual appearance of the intraoperative scene in the context of their prior knowledge. Vision-based navigation systems seek to replicate this capability by recovering camera pose directly from endoscopic video, but most approaches do not embody the same principles of reasoning about new frames that makes surgeons successful. Instead, they remain grounded in feature matching and geometric optimization over keyframes, an approach that has been shown to degrade under the challenging conditions of endoscopic imaging like low texture and rapid illumination changes. Here, we pursue an alternative approach and investigate a policy-based formulation of endoscopic camera pose recovery that seeks to imitate experts in estimating trajectories conditioned on the previous camera state. Our approach directly predicts short-horizon relative motions without maintaining an explicit geometric representation at inference time. It thus addresses, by design, some of the notorious challenges of geometry-based approaches, such as brittle correspondence matching, instability in texture-sparse regions, and limited pose coverage due to reconstruction failure. We evaluate the proposed formulation on cadaveric sinus endoscopy. Under oracle state conditioning, we compare short-horizon motion prediction quality to geometric baselines achieving lowest mean translation error and competitive rotational accuracy. We analyze robustness by grouping prediction windows according to texture richness and illumination change indicating reduced sensitivity to low-texture conditions. These findings suggest that a learned motion policy offers a viable alternative formulation for endoscopic camera pose recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。