梳理大模型在牙科中的应用,揭示三类模型的优劣与融合路径。
Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models

- 按架构与牙科专业化程度构建二维分类框架
- 牙科专用模型在多模态任务中表现最佳,整合管道优于单一模型
- 适用于牙科AI研究者、临床开发者及政策制定者参考
口腔疾病影响全球近35亿人,但大规模AI模型在牙科中的临床潜力尚不明确。目前出现三类模型:语言生成模型、判别性视觉基础模型和牙科专用基础模型,缺乏统一综述分析其关系与共同局限。根据PRISMA-ScR指南,系统检索四个数据库(PubMed、Google Scholar、Scopus、arXiv),两名评审员独立筛选,最终纳入97项研究(2020–2026)。提出二维分类框架,按架构范式与牙科专业化程度划分。语言生成模型在文本任务(临床推理、执照考试、患者沟通)中表现优异,但在图像诊断上表现不一;改进版SAM与CLIP在牙齿分割与病灶检测中效果良好。牙科专用模型(DentVFM、DentVLM、OralGPT)在复杂多模态任务中表现最强;集成式流程持续优于单一模型。数据存在不对称:牙科预训练几乎全部集中于视觉领域,反映大规模牙科文本语料稀缺。结论:通用模型与专用模型互补,最有效系统需在结构化流程中融合两者。安全自主部署仍需突破三大障碍:生成模型幻觉、标注数据少、缺乏标准化临床评估基准。
原文摘要 · Abstract (English)
Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical potential of large-scale AI models in dentistry remains poorly understood. Three distinct model categories have emerged: language-generative models, discriminative vision foundation models, and dental-specific foundation models, with no unified review examining their relationships and collective limitations. Methods: Following PRISMA-ScR guidelines, we systematically searched four databases (PubMed, Google Scholar, Scopus, arXiv), screened independently by two reviewers. After applying inclusion/exclusion criteria, 97 studies (2020-2026) were included. We propose a two-dimensional classification framework organizing models by architectural paradigm and dental specialization degree. Results: Language-generative models excel at text-based tasks (clinical reasoning, licensing exams, patient communication) but show inconsistent performance on image-dependent diagnostics. Adapted SAM and CLIP variants achieve strong tooth segmentation and lesion detection results. Dental-specific models (DentVFM, DentVLM, OralGPT) demonstrate strongest performance on complex multimodal tasks. Integrated pipelines consistently outperform single-model approaches. A data asymmetry is observed: dental-specific pretraining concentrates almost entirely in the vision domain, reflecting scarce large-scale dental text corpora. Conclusions: General-purpose and dental-specific models play complementary roles; the most effective systems combine both within structured pipelines. Safe autonomous deployment requires resolving three persistent barriers: hallucination in generative models, limited annotated dental datasets, and absent standardized clinical evaluation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。