对比深度学习与影像组学在胸部X光诊断中的表现,发现数据多时深度学习更优。
Comparative Evaluation of Radiomics and Deep Learning Models for Disease Detection in Chest Radiography
- 比较影像组学特征提取与深度学习直接图像分析的诊断效果
- 4000样本下InceptionV3 AUC达0.996,显著优于传统模型
- 数据少时影像组学仍可用,大数据下深度学习更具优势
人工智能在医学影像中的应用已革新诊断流程,本研究系统评估了基于影像组学和深度学习的胸部X光疾病检测方法,聚焦新冠、肺部阴影和病毒性肺炎。影像组学模型采用手工特征提取,深度学习模型如卷积神经网络和视觉转换器则直接从图像学习。对比了决策树、梯度提升、随机森林、支持向量机和多层感知机等影像组学模型,以及InceptionV3、EfficientNetL和ConvNeXtXLarge等先进深度学习模型。在24个样本下,EfficientNetL AUC为0.839,优于SVM的0.762;在4000样本下,InceptionV3 AUC达0.996,远超随机森林的0.885。方差分析显示模型类型和样本量对所有指标均有显著主效应和交互效应。事后检验确认深度学习在多数条件下表现更优。研究为不同数据规模下的AI模型选型提供数据驱动建议:深度学习在数据充足时性能更强且可扩展,影像组学适用于数据受限场景。
原文摘要 · Abstract (English)
The application of artificial intelligence (AI) in medical imaging has revolutionized diagnostic practices, enabling advanced analysis and interpretation of radiological data. This study presents a comprehensive evaluation of radiomics-based and deep learning-based approaches for disease detection in chest radiography, focusing on COVID-19, lung opacity, and viral pneumonia. While deep learning models, particularly convolutional neural networks and vision transformers, learn directly from image data, radiomics-based models extract handcrafted features, offering potential advantages in data-limited scenarios. We systematically compared the diagnostic performance of various AI models, including Decision Trees, Gradient Boosting, Random Forests, Support Vector Machines, and Multi-Layer Perceptrons for radiomics, against state-of-the-art deep learning models such as InceptionV3, EfficientNetL, and ConvNeXtXLarge. Performance was evaluated across multiple sample sizes. At 24 samples, EfficientNetL achieved an AUC of 0.839, outperforming SVM (AUC = 0.762). At 4000 samples, InceptionV3 achieved the highest AUC of 0.996, compared to 0.885 for Random Forest. A Scheirer-Ray-Hare test confirmed significant main and interaction effects of model type and sample size on all metrics. Post hoc Mann-Whitney U tests with Bonferroni correction further revealed consistent performance advantages for deep learning models across most conditions. These findings provide statistically validated, data-driven recommendations for model selection in diagnostic AI. Deep learning models demonstrated higher performance and better scalability with increasing data availability, while radiomics-based models may remain useful in low-data contexts. This study addresses a critical gap in AI-based diagnostic research by offering practical guidance for deploying AI models across diverse clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。