用扩散模型生成带标注的头影测量图像,提升牙科检测精度
Towards Better Cephalometric Landmark Detection with Diffusion Data Generation
- 基于解剖先验构建新标注,结合扩散模型生成真实头影测量片
- 生成数据训练下成功检测率提升至82.2%,较基线高6.5%
- 适合需要高质量医学图像数据的医疗AI研究者使用
头影测量点检测对正畸诊断与治疗规划至关重要,但数据采集样本少、人工标注耗时,严重制约了多样化数据集的获取,限制了基于深度学习的检测方法,尤其是大规模视觉模型的应用。为此,我们提出一种创新的数据生成方法,可无需人工干预地生成多样化的头影测量X光片及其对应标注。该方法首先利用解剖先验构建新的头影测量点标注,再通过扩散生成器生成与标注高度匹配的逼真X光图像。为实现对不同属性样本的精确控制,我们构建了一个包含真实头影测量图像和详细医学文本描述的提示数据集。借助该数据集,方法能有效控制生成图像的风格与特征。在大量多样化的生成数据支持下,我们将大规模视觉检测模型引入头影测量点检测任务,显著提升准确率。实验表明,使用生成数据训练后,成功检测率(SDR)提升6.5%,达到82.2%。所有代码与数据已公开:https://um-lab.github.io/cepha-generation
原文摘要 · Abstract (English)
Cephalometric landmark detection is essential for orthodontic diagnostics and treatment planning. Nevertheless, the scarcity of samples in data collection and the extensive effort required for manual annotation have significantly impeded the availability of diverse datasets. This limitation has restricted the effectiveness of deep learning-based detection methods, particularly those based on large-scale vision models. To address these challenges, we have developed an innovative data generation method capable of producing diverse cephalometric X-ray images along with corresponding annotations without human intervention. To achieve this, our approach initiates by constructing new cephalometric landmark annotations using anatomical priors. Then, we employ a diffusion-based generator to create realistic X-ray images that correspond closely with these annotations. To achieve precise control in producing samples with different attributes, we introduce a novel prompt cephalometric X-ray image dataset. This dataset includes real cephalometric X-ray images and detailed medical text prompts describing the images. By leveraging these detailed prompts, our method improves the generation process to control different styles and attributes. Facilitated by the large, diverse generated data, we introduce large-scale vision detection models into the cephalometric landmark detection task to improve accuracy. Experimental results demonstrate that training with the generated data substantially enhances the performance. Compared to methods without using the generated data, our approach improves the Success Detection Rate (SDR) by 6.5%, attaining a notable 82.2%. All code and data are available at: https://um-lab.github.io/cepha-generation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。