arXiv:2509.09911cs.CVcs.AI2025-09

用自编码器+视觉Transformer提升牙齿年龄判读准确率并解释模型不确定性

An Autoencoder and Vision Transformer-based Interpretability Analysis of the Differences in Automated Staging of Second and Third Molars

  • 融合自编码器与视觉Transformer,提升下颌第二、第三磨牙自动分期准确率
  • 第二磨牙准确率从0.712升至0.815,第三磨牙从0.462升至0.543
  • 通过潜空间分析发现第三磨牙数据内部差异大是性能瓶颈,适合法医鉴定场景

深度学习在高风险法医应用(如牙齿年龄估计)中的实际应用常受限于模型的‘黑箱’特性。本研究以下颌第二磨牙(牙37)和第三磨牙(牙38)自动分期表现差异为案例,提出一种结合卷积自编码器(AE)与视觉Transformer(ViT)的框架。该框架在基准ViT基础上提升了两类牙齿的分类准确率:牙37从0.712提升至0.815,牙38从0.462提升至0.543。除性能提升外,框架还提供多维度诊断洞察。自编码器潜空间指标与图像重建分析表明,性能差距主要源于数据问题,牙38数据集内部形态变异性高是关键制约因素。研究强调仅依赖注意力图等单一可解释性手段可能误导判断,无法揭示底层数据缺陷。该框架既提升准确率,又提供模型不确定性的实证依据,为法医年龄估计中的专家决策提供更可靠的支撑。

原文摘要 · Abstract (English)

The practical adoption of deep learning in high-stakes forensic applications, such as dental age estimation, is often limited by the 'black box' nature of the models. This study introduces a framework designed to enhance both performance and transparency in this context. We use a notable performance disparity in the automated staging of mandibular second (tooth 37) and third (tooth 38) molars as a case study. The proposed framework, which combines a convolutional autoencoder (AE) with a Vision Transformer (ViT), improves classification accuracy for both teeth over a baseline ViT, increasing from 0.712 to 0.815 for tooth 37 and from 0.462 to 0.543 for tooth 38. Beyond improving performance, the framework provides multi-faceted diagnostic insights. Analysis of the AE's latent space metrics and image reconstructions indicates that the remaining performance gap is data-centric, suggesting high intra-class morphological variability in the tooth 38 dataset is a primary limiting factor. This work highlights the insufficiency of relying on a single mode of interpretability, such as attention maps, which can appear anatomically plausible yet fail to identify underlying data issues. By offering a methodology that both enhances accuracy and provides evidence for why a model may be uncertain, this framework serves as a more robust tool to support expert decision-making in forensic age estimation.

牙齿年龄估计可解释性视觉Transformer自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。