直接用3D点云做口腔扫描诊断,让AI能统一识别23种牙病并回答问题。
IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans
- 将口腔扫描转为点云,用3D编码器+大模型实现端到端诊断
- 在1.9万病例上构建了大规模多源VQA数据集,覆盖23种疾病
- 通过几何转色彩代理提升精度,适合临床医生和牙科研究者
3D口腔扫描因丰富的几何信息正广泛用于临床,统一多病诊断对病历记录与沟通至关重要。现有工作虽引入牙科视觉语言模型(VLM)对2D图像或多视角图像进行诊断与报告生成,但未充分挖掘原生3D几何信息。该任务面临三大挑战:(i) 扫描形式多样且拓扑复杂,(ii) 多病共现、类别不平衡及细微形态模糊,(iii) 配对3D扫描-文本数据有限。为此,我们提出IOSVLM,一种端到端3D VLM,将扫描表示为点云,采用3D编码器-投影器-大语言模型架构,实现统一诊断与生成式视觉问答(VQA)。同时构建了大型多源数据集IOSVQA,包含19,002个病例与249,055个VQA对,覆盖23种口腔疾病和异构扫描类型。为缓解无色扫描数据与依赖颜色的3D预训练之间的分布差距,提出几何转色彩代理机制,稳定细粒度几何感知与跨模态对齐。采用两阶段课程训练策略进一步提升鲁棒性。实验表明,IOSVLM显著优于强基线,宏平均准确率提升至少+9.58%,宏平均F1提升+1.46%,验证了直接建模3D几何信息在基于3D扫描诊断中的有效性。
原文摘要 · Abstract (English)
3D intraoral scans (IOS) are increasingly adopted in routine dentistry due to abundant geometric evidence, and unified multi-disease diagnosis is desirable for clinical documentation and communication. While recent works introduce dental vision-language models (VLMs) to enable unified diagnosis and report generation on 2D images or multi-view images rendered from IOS, they do not fully leverage native 3D geometry. Such work is necessary and also challenging, due to: (i) heterogeneous scan forms and the complex IOS topology, (ii) multi-disease co-occurrence with class imbalance and fine-grained morphological ambiguity, (iii) limited paired 3D IOS-text data. Thus, we present IOSVLM, an end-to-end 3D VLM that represents scans as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative visual question-answering (VQA), together with IOSVQA, a large-scale multi-source IOS diagnosis VQA dataset comprising 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types. To address the distribution gap between color-free IOS data and color-dependent 3D pre-training, we propose a geometry-to-chromatic proxy that stabilizes fine-grained geometric perception and cross-modal alignment. A two-stage curriculum training strategy further enhances robustness. IOSVLM consistently outperforms strong baselines, achieving gains of at least +9.58% macro accuracy and +1.46% macro F1, indicating the effectiveness of direct 3D geometry modeling for IOS-based diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。