DualTrack用双路架构分离处理局部与全局信息,实现高精度无传感器3D超声重建。
DualTrack: Sensorless 3D Ultrasound needs Local and Global Context
- 采用双编码器分离提取帧间细节与整体解剖结构特征
- 在公开数据集上平均重建误差低于5毫米,性能领先现有方法
- 适合需要低成本3D超声成像的临床场景与科研应用
三维超声(US)相比传统2D成像具有诸多临床优势,但其广泛应用受限于传统3D系统的成本与复杂性。无传感器3D超声通过深度学习从2D图像序列中估计探头轨迹,是一种有前景的替代方案。局部特征(如斑点纹理)可预测帧间运动,而全局特征(如粗略形态与解剖结构)能定位扫描相对于解剖的位置并预测整体形状。以往方法要么忽略全局特征,要么将其与局部特征提取紧密耦合,限制了对这两种互补信息的鲁棒建模。我们提出DualTrack,一种新型双编码器架构,通过解耦的局部与全局编码器分别专注于各自尺度的特征提取。局部编码器使用密集时空卷积捕捉细粒度特征,全局编码器则采用图像主干网络(如2D CNN或基础模型)结合时间注意力层,嵌入高层解剖特征与长程依赖关系。一个轻量级融合模块将两者特征结合以估计轨迹。在大型公开基准上的实验结果表明,DualTrack达到当前最优精度,生成全局一致的3D重建,优于先前方法,平均重建误差低于5毫米。
原文摘要 · Abstract (English)
Three-dimensional ultrasound (US) offers many clinical advantages over conventional 2D imaging, yet its widespread adoption is limited by the cost and complexity of traditional 3D systems. Sensorless 3D US, which uses deep learning to estimate a 3D probe trajectory from a sequence of 2D US images, is a promising alternative. Local features, such as speckle patterns, can help predict frame-to-frame motion, while global features, such as coarse shapes and anatomical structures, can situate the scan relative to anatomy and help predict its general shape. In prior approaches, global features are either ignored or tightly coupled with local feature extraction, restricting the ability to robustly model these two complementary aspects. We propose DualTrack, a novel dual-encoder architecture that leverages decoupled local and global encoders specialized for their respective scales of feature extraction. The local encoder uses dense spatiotemporal convolutions to capture fine-grained features, while the global encoder utilizes an image backbone (e.g., a 2D CNN or foundation model) and temporal attention layers to embed high-level anatomical features and long-range dependencies. A lightweight fusion module then combines these features to estimate the trajectory. Experimental results on a large public benchmark show that DualTrack achieves state-of-the-art accuracy and globally consistent 3D reconstructions, outperforming previous methods and yielding an average reconstruction error below 5 mm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。