VerteNet精准定位侧位脊柱DXA图像中的椎体关键点,提升骨折评估与手术导航精度。
VerteNet -- A Multi-Context Hybrid CNN Transformer for Accurate Vertebral Landmark Localization in Lateral Spine DXA Images
- 融合双分辨率注意力机制,兼顾局部细节与全局上下文信息
- 在多设备数据上实现4.92的归一化均值误差和2.35的中位误差
- 适用于低信噪比图像,适合临床脊柱评估与手术规划场景
双能X射线吸收测定术(DXA)侧位脊柱成像中的椎体关键点定位对评估脊柱对齐、椎体骨折及腹主动脉钙化量化时的椎间引导放置至关重要。尽管侧位脊柱DXA扫描具有成本低、辐射少的优势,但其分析因信噪比低和成像伪影而困难。人工智能为提升椎体关键点定位的精度提供了可行方案。本文提出一种新型混合架构VerteNet,采用双分辨率注意力机制,同时捕捉精细局部细节与广泛上下文信息。通过跳跃连接和解码器层结合双分辨率自注意力与交叉注意力机制,增强特征融合能力,使模型更有效地学习复杂模式,实现精确的椎体角点定位,同时保持局部与全局感知。我们在多个设备采集的DXA LSI图像上评估该框架,结果表明其性能优于现有先进方法,归一化均值误差为4.92,归一化中位误差为2.35。VerteNet在不同采集系统下均实现高精度关键点定位,并展现出对低信噪比图像的强鲁棒性,归功于其对细粒度局部细节与广义上下文信息的双重捕捉能力。
原文摘要 · Abstract (English)
Vertebral Landmarks Localization in Dual-Energy X-ray Absorptiometry based Lateral Spine Imaging plays a critical role in evaluating spinal alignment, Vertebral Fracture Assessment, and facilitating intervertebral guide placement for Abdominal Aortic Calcification quantification. While lateral spine DXA scans offer advantages such as reduced cost and lower radiation exposure, its analysis remains challenging due to a low signal-to-noise ratio and imaging artifacts. Artificial Intelligence presents a promising approach for improving the precision and accuracy of VLL. In this study, we introduce a novel architecture that employs dual-resolution attention mechanisms to capture both fine-grained local details and broader contextual information. Our approach enhances feature integration by leveraging skip connections and decoder layers through dual-resolution self-attention and cross-attention mechanisms. This design improves the ability of the model to learn complex patterns, enabling precise vertebral corner localization while maintaining both local and global contextual awareness. We evaluated the proposed framework on DXA LSI images acquired from multiple machines and found that it outperforms recent state-of-the-art architectures for VLL, achieving a normalized mean error of 4.92 and a normalized median error of 2.35. The proposed framework, VerteNet, enables highly accurate VLL in DXA LSI images from diverse acquisition systems and demonstrates strong robustness to low signal-to-noise ratios, owing to its enhanced ability to capture both fine-grained local details and broader contextual information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。