arXiv:2607.12398cs.CVcs.RO2026-07

用全局相机+局部超声信号,提升自由手3D超声重建精度

Seeing Globally, Refining Locally: Global Visual Guidance and Local Ultrasound Cues for Robust Freehand 3-D Ultrasound Reconstruction

论文配图:Seeing Globally, Refining Locally: Global Visual Guidance and Local Ultrasound Cues for Robust Freehand 3-D Ultrasound Reconstruction
图 1 · 摘自论文原文
  • 通过双摄像头和超声图像联合建模,实现全局稳定定位与局部组织自适应修正
  • 在真实人体数据上轨迹误差低至1.29毫米,比纯超声方法减少超过13毫米
  • 适合临床场景中长时间、高精度的自由手超声重建任务

自由手3D超声成像因其直观的体积可视化、易用性和低成本而受到关注。然而,精确重建高度依赖稳定的探头位姿估计,现有无追踪器方法在长扫描路径下仍易受累积位姿误差影响。为此,我们提出一种从全局到局部的位姿估计框架,利用外部相机观测实现全局稳定定位,并结合B模式超声图像进行解剖感知的局部精修。该框架包含双摄像头分支,通过跨视角和时序特征聚合估计全局一致的探头轨迹;以及超声分支,通过序列超声图像中的解剖特征聚合捕捉组织相关的局部运动线索。一个跨模态融合模块将相机上下文特征与超声解剖特征融合,预测位姿残差并在变换空间中修正相机估计结果。此外,多尺度位姿损失在多个时间跨度上约束相对运动,抑制长时间扫描中的漂移。在模拟物和活体数据集上验证表明,在两个自建数据集(FUSION-J 和 FUSION-L)上,所提的US + Dual-Cam模型将平均轨迹漂移分别降低至1.67毫米和1.29毫米,相比强基线双摄像头方法分别提升16.50%和27.12%,显著优于仅使用超声的方法(>13毫米漂移)。在活体前臂动脉重建中,达到1.58毫米的豪斯多夫距离,证明了该方法在真实临床场景中的有效性。

原文摘要 · Abstract (English)

Freehand 3-D ultrasound (US) imaging has attracted increasing attention owing to its intuitive volumetric visualization, ease of use, and low cost. However, accurate 3-D reconstruction critically depends on stable probe pose estimation, yet existing trackerless methods remain susceptible to accumulated pose errors, particularly over long scanning trajectories. To address this limitation, we propose a global-to-local pose estimation framework that exploits external camera observations for globally stable localization and B-mode US images for anatomy-aware local refinement. Specifically, the framework comprises a dual-camera branch that performs contextual feature aggregation across camera views and temporal observations to estimate a globally consistent probe trajectory, and a B-mode branch that performs anatomical feature aggregation from sequential US images to capture tissue-dependent local motion cues. A cross-modal fusion module subsequently integrates the contextual camera features and anatomical US features to predict pose residuals and refine the camera-derived estimates in the transformation space. Furthermore, a multi-scale pose loss constrains relative motion over multiple temporal horizons to suppress accumulated drift during extended scans. The proposed framework is validated on phantom and in vivo datasets. On two in-house datasets (FUSION-J and FUSION-L) collected using different machines, the proposed US + Dual-Cam model reduces average trajectory drift to 1.67 mm and 1.29 mm, representing improvement of 16.50% and 27.12%, respectively, over a strong dual-camera baseline, while substantially outperforming US-only pose estimation (>13 mm drift). In in vivo forearm arteries reconstruction, it achieves Hausdorff distances of 1.58 mm, demonstrating the effectiveness of the proposed method on real clinical scenarios.

3D超声位姿估计跨模态融合医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。