arXiv:2603.17178cs.CV2026-03被引 1

在手术室单目视频中,稳定重建患者3D身体网格。

Patient4D: Temporally Consistent Patient Body Mesh Recovery from Monocular Operating Room Video

  • 利用静止先验和关键帧锚定姿态,提升时序一致性。
  • 遮挡下平均交并比达0.75,失败帧从30.5%降至1.3%。
  • 适合临床增强现实中的患者体表动态重建场景。

从单目手术室视频中恢复密集3D身体网格仍面临布帘遮挡与摄像机视角持续变化的挑战。该场景出现在手术增强现实(AR)中,麻醉患者被手术布帘覆盖,而外科医生头戴摄像头导致视角不断变化。现有人体网格重建(HMR)方法通常在直立运动主体、相对稳定的摄像机下训练,因此在此类条件下性能下降。为此,我们提出Patient4D,一种基于静止先验的重建流程,结合图像级基础模型与轻量几何机制,实现跨帧时序一致性。两个关键组件提升鲁棒性:姿态锁定(Pose Locking)通过稳定关键帧锚定姿态参数;刚性回退(Rigid Fallback)在严重遮挡下通过轮廓引导刚性对齐恢复网格。二者协同稳定预测,且兼容现成的HMR模型。我们在4,680个合成手术序列及三个公开HMR视频基准上评估。在手术布帘遮挡下,Patient4D达到0.75平均交并比(IoU),将失败帧率从最佳基线的30.5%降至1.3%。结果表明,利用静止先验可显著提升临床AR场景下的单目重建效果。

原文摘要 · Abstract (English)

Recovering a dense 3D body mesh from monocular video remains challenging under occlusion from draping and continuously moving camera viewpoints. This configuration arises in surgical augmented reality (AR), where an anesthetized patient lies under surgical draping while a surgeon's head-mounted camera continuously changes viewpoint. Existing human mesh recovery (HMR) methods are typically trained on upright, moving subjects captured from relatively stable cameras, leading to performance degradation under such conditions. To address this, we present Patient4D, a stationarity-constrained reconstruction pipeline that explicitly exploits the stationarity prior. The pipeline combines image-level foundation models for perception with lightweight geometric mechanisms that enforce temporal consistency across frames. Two key components enable robust reconstruction: Pose Locking, which anchors pose parameters using stable keyframes, and Rigid Fallback, which recovers meshes under severe occlusion through silhouette-guided rigid alignment. Together, these mechanisms stabilize predictions while remaining compatible with off-the-shelf HMR models. We evaluate Patient4D on 4,680 synthetic surgical sequences and three public HMR video benchmarks. Under surgical drape occlusion, Patient4D achieves a 0.75 mean IoU, reducing failure frames from 30.5% to 1.3% compared to the best baseline. Our findings demonstrate that exploiting stationarity priors can substantially improve monocular reconstruction in clinical AR scenarios.

3D重建手术视觉时序一致单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。