arXiv:2511.14336cs.CV2025-11被引 4

无需训练,用几何标准化+知识库实现牙齿计数与结构解析

ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding

  • 通过几何感知的牙弓展平技术统一3D扫描形态
  • 在1060例病例中实现98.7%牙齿计数准确率,抗遮挡和噪声能力强
  • 适合临床部署,尤其适用于设备多样、标注少的真实场景

口腔内3D扫描的结构化理解对数字化正畸至关重要。现有深度学习方法依赖特定模态训练、大规模标注数据及受控扫描条件,限制了跨设备泛化能力,难以融入真实临床流程。原始牙弓网格存在姿态差异大、因咬合或牙齿接触导致的几何不完整,且缺乏纹理信息,统一语义解析难度高。为此,我们提出ArchMap——一种无需训练、基于知识引导的鲁棒结构化牙科理解框架。首先引入几何感知牙弓展平模块,将原始3D网格转换为空间对齐、连续性保持的多视角投影;随后构建牙科知识库(DKB),编码层级牙齿本体、萌出阶段策略与临床语义,约束符号推理空间。在1060例正畸前后病例上验证,ArchMap在牙齿计数、解剖分区、萌出阶段分类以及拥挤、缺牙、义齿、龋齿等临床状况识别上表现优异。相比监督流水线与提示式视觉语言模型基线,其精度更高,语义漂移更小,在稀疏或含伪影条件下稳定性更强。作为全免训练系统,ArchMap证明了几何归一化与本体引导多模态推理结合,是现代数字正畸中3D口腔扫描结构化分析的实用且可扩展方案。

原文摘要 · Abstract (English)

A structured understanding of intraoral 3D scans is essential for digital orthodontics. However, existing deep-learning approaches rely heavily on modality-specific training, large annotated datasets, and controlled scanning conditions, which limit generalization across devices and hinder deployment in real clinical workflows. Moreover, raw intraoral meshes exhibit substantial variation in arch pose, incomplete geometry caused by occlusion or tooth contact, and a lack of texture cues, making unified semantic interpretation highly challenging. To address these limitations, we propose ArchMap, a training-free and knowledge-guided framework for robust structured dental understanding. ArchMap first introduces a geometry-aware arch-flattening module that standardizes raw 3D meshes into spatially aligned, continuity-preserving multi-view projections. We then construct a Dental Knowledge Base (DKB) encoding hierarchical tooth ontology, dentition-stage policies, and clinical semantics to constrain the symbolic reasoning space. We validate ArchMap on 1060 pre-/post-orthodontic cases, demonstrating robust performance in tooth counting, anatomical partitioning, dentition-stage classification, and the identification of clinical conditions such as crowding, missing teeth, prosthetics, and caries. Compared with supervised pipelines and prompted VLM baselines, ArchMap achieves higher accuracy, reduced semantic drift, and superior stability under sparse or artifact-prone conditions. As a fully training-free system, ArchMap demonstrates that combining geometric normalization with ontology-guided multimodal reasoning offers a practical and scalable solution for the structured analysis of 3D intraoral scans in modern digital orthodontics.

牙齿计数3D扫描知识图谱数字正畸

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。