用一张全景牙片重建3D口腔结构,无需额外扫描或先验信息。
ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs
- 融合视觉Transformer与神经啤酒-兰伯特模型,直接从单张牙片重建3D结构。
- 新马蹄形采样策略减少52%采样计算量,避免射线交叉导致的密度聚合。
- 自学习哈希位置编码提升3D点表示能力,适合临床无先验数据场景。
牙科诊断依赖两种主要影像方式:全景牙片(PX)提供二维口腔图像,锥形束CT(CBCT)则能获取详细的三维解剖信息。虽然PX成本低、易获取,但缺乏深度信息影响诊断准确性;而CBCT虽可弥补,却存在成本高、辐射大、普及难等问题。现有重建方法常需对CBCT进行展平处理或依赖已知牙弓信息,临床中往往不可得。本文提出ViT-NeBLa,一种基于视觉变压器的神经啤酒-兰伯特框架,实现仅凭单张全景牙片完成精确3D重建。核心创新包括:(1) 将视觉变压器融入NeBLa框架,无需CBCT展平或牙弓先验信息即可增强重建性能;(2) 设计新型马蹄形点采样策略,采用非交叉射线路径,消除传统方法中因射线交叉引起的中间密度聚合,计算量降低52%;(3) 以混合视觉变压器-卷积网络替代原有基于CNN的U-Net,实现更优的全局与局部特征提取;(4) 引入可学习哈希位置编码,相比现有基于傅里叶的密集位置编码,更优地表示高维3D采样点。实验表明,ViT-NeBLa在定量和定性上均显著优于现有最先进方法,为牙科诊断提供低成本、低辐射的高效解决方案。
原文摘要 · Abstract (English)
Dental diagnosis relies on two primary imaging modalities: panoramic radiographs (PX) providing 2D oral cavity representations, and Cone-Beam Computed Tomography (CBCT) offering detailed 3D anatomical information. While PX images are cost-effective and accessible, their lack of depth information limits diagnostic accuracy. CBCT addresses this but presents drawbacks including higher costs, increased radiation exposure, and limited accessibility. Existing reconstruction models further complicate the process by requiring CBCT flattening or prior dental arch information, often unavailable clinically. We introduce ViT-NeBLa, a vision transformer-based Neural Beer-Lambert model enabling accurate 3D reconstruction directly from single PX. Our key innovations include: (1) enhancing the NeBLa framework with Vision Transformers for improved reconstruction capabilities without requiring CBCT flattening or prior dental arch information, (2) implementing a novel horseshoe-shaped point sampling strategy with non-intersecting rays that eliminates intermediate density aggregation required by existing models due to intersecting rays, reducing sampling point computations by $52 \%$, (3) replacing CNN-based U-Net with a hybrid ViT-CNN architecture for superior global and local feature extraction, and (4) implementing learnable hash positional encoding for better higher-dimensional representation of 3D sample points compared to existing Fourier-based dense positional encoding. Experiments demonstrate that ViT-NeBLa significantly outperforms prior state-of-the-art methods both quantitatively and qualitatively, offering a cost-effective, radiation-efficient alternative for enhanced dental diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。