用面部图像诊断库欣综合征,视觉变压器模型表现优于传统卷积网络。
Comparative Analysis of Pre-trained Deep Learning Models and DINOv2 for Cushing's Syndrome Diagnosis in Facial Analysis
- 采用视觉变压器和DINOv2模型捕捉面部全局特征,提升诊断精度。
- ViT模型达到85.74%的最高F1分数,女性样本识别更准确。
- 冻结参数后DINOv2性能提升,适合医疗影像少样本场景使用。
库欣综合征由肾上腺皮质过度分泌糖皮质激素引起,常表现为满月脸和面色潮红,面部数据对诊断至关重要。以往研究多使用预训练卷积神经网络(CNN)分析正面面部图像进行诊断,但CNN擅长局部特征,而库欣综合征更多体现为全局面部特征。基于Transformer的ViT和SWIN等模型利用自注意力机制,更利于捕捉长程依赖与整体特征。近期,基于视觉Transformer的DINOv2基础模型受到关注。本研究对比了多种预训练模型(包括CNN、Transformer模型及DINOv2)在库欣综合征诊断中的表现,并分析了性别偏差及参数冻结机制对DINOv2的影响。结果表明,基于Transformer的模型与DINOv2均优于CNN,其中ViT取得85.74%的最高F1分数;预训练模型和DINOv2对女性样本的准确率更高;冻结参数后DINOv2性能进一步提升。结论:基于Transformer的模型和DINOv2在库欣综合征分类中效果显著。
原文摘要 · Abstract (English)
Cushing's syndrome is a condition caused by excessive glucocorticoid secretion from the adrenal cortex, often manifesting with moon facies and plethora, making facial data crucial for diagnosis. Previous studies have used pre-trained convolutional neural networks (CNNs) for diagnosing Cushing's syndrome using frontal facial images. However, CNNs are better at capturing local features, while Cushing's syndrome often presents with global facial features. Transformer-based models like ViT and SWIN, which utilize self-attention mechanisms, can better capture long-range dependencies and global features. Recently, DINOv2, a foundation model based on visual Transformers, has gained interest. This study compares the performance of various pre-trained models, including CNNs, Transformer-based models, and DINOv2, in diagnosing Cushing's syndrome. We also analyze gender bias and the impact of freezing mechanisms on DINOv2. Our results show that Transformer-based models and DINOv2 outperformed CNNs, with ViT achieving the highest F1 score of 85.74%. Both the pre-trained model and DINOv2 had higher accuracy for female samples. DINOv2 also showed improved performance when freezing parameters. In conclusion, Transformer-based models and DINOv2 are effective for Cushing's syndrome classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。