arXiv:2509.23100cs.CV2025-09

对比五种视觉Transformer模型在牙科疾病分类中的表现,发现ConvNeXt最准确。

Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer

  • 用五种主流Transformer架构对比牙科影像分类性能
  • ConvNeXt准确率达81.06,优于其他模型
  • 强调数据不平衡对模型的影响,适合临床诊断应用

本研究系统评估并比较了五种前沿基于Transformer的图像模型——视觉变换器(ViT)、数据高效图像变换器(DeiT)、ConvNeXt、Swin Transformer和双向图像变换器编码器(BEiT)在多类牙病分类中的表现。研究聚焦真实世界挑战,如数据不平衡问题,该问题常被现有文献忽视。采用口腔疾病数据集训练与验证所选模型,以验证准确率、精确率、召回率和F1分数为评估指标,特别关注各架构在类别不平衡情况下的表现。结果显示,ConvNeXt达到最高验证准确率81.06,其次为BEiT(80.00)和Swin Transformer(79.73),三者均表现出优异的F1分数。ViT和DeiT分别取得79.37和78.79的准确率,但在龋齿相关类别上表现较差。结论表明,ConvNeXt、Swin Transformer和BEiT展现出可靠的诊断性能,具备临床应用潜力,为未来基于AI的牙科疾病诊断工具提供模型选择依据,并强调在真实场景中应对数据不平衡的重要性。

原文摘要 · Abstract (English)

Objective: The aim of this study was to systematically evaluate and compare the performance of five state-of-the-art transformer-based architectures - Vision Transformer (ViT), Data-efficient Image Transformer (DeiT), ConvNeXt, Swin Transformer, and Bidirectional Encoder Representation from Image Transformers (BEiT) - for multi-class dental disease classification. The study specifically focused on addressing real-world challenges such as data imbalance, which is often overlooked in existing literature. Study Design: The Oral Diseases dataset was used to train and validate the selected models. Performance metrics, including validation accuracy, precision, recall, and F1-score, were measured, with special emphasis on how well each architecture managed imbalanced classes. Results: ConvNeXt achieved the highest validation accuracy at 81.06, followed by BEiT at 80.00 and Swin Transformer at 79.73, all demonstrating strong F1-scores. ViT and DeiT achieved accuracies of 79.37 and 78.79, respectively, but both struggled particularly with Caries-related classes. Conclusions: ConvNeXt, Swin Transformer, and BEiT showed reliable diagnostic performance, making them promising candidates for clinical application in dental imaging. These findings provide guidance for model selection in future AI-driven oral disease diagnostic tools and highlight the importance of addressing data imbalance in real-world scenarios

牙科影像Transformer分类模型数据不平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。