综述多模态深度学习在新冠诊断中的应用,揭示图像、文本、语音的高精度识别潜力。
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
- 整合图像、文本、语音多模态数据,构建11种深度学习模型进行对比分析
- MobileNet在图像和语音识别中分别达99.97%和93.73%准确率,BiGRU在文本分类中达99.89%
- 为医疗AI跨模态诊断提供方法参考,适合医学与人工智能交叉研究者
本研究系统综述了多模态深度学习(DL)在医学诊断中的潜力,以新冠检测为例。基于AI在新冠疫情中的成功应用,研究旨在揭示深度学习在疾病筛查、预测与分类中的能力,并为科技与创新体系的韧性、可持续性与包容性提供洞见。通过系统分析不同研究的方法、数据来源、预处理流程与挑战,探讨了深度学习模型的架构设计及其数据适配性。进一步比较了多种深度学习策略在新冠分析中的表现,评估其方法、数据、性能及未来研究前提。通过对多类型数据与诊断模态的考察,本研究推动了对多模态深度学习应用于诊断的有效性理解。我们基于新冠图像、文本与语音(即咳嗽)数据,实现并分析了11种深度学习模型。结果表明,MobileNet在图像数据上达到99.97%准确率,在语音数据上达93.73%;而BiGRU在文本分类中表现最优,准确率达99.89%。研究结果对其他需图像、文本、语音分析的领域具有潜在借鉴价值。
原文摘要 · Abstract (English)
This study presents a comprehensive review of the potential of multimodal deep learning (DL) in medical diagnosis, using COVID-19 as a case example. Motivated by the success of artificial intelligence applications during the COVID-19 pandemic, this research aims to uncover the capabilities of DL in disease screening, prediction, and classification, and to derive insights that enhance the resilience, sustainability, and inclusiveness of science, technology, and innovation systems. Adopting a systematic approach, we investigate the fundamental methodologies, data sources, preprocessing steps, and challenges encountered in various studies and implementations. We explore the architecture of deep learning models, emphasising their data-specific structures and underlying algorithms. Subsequently, we compare different deep learning strategies utilised in COVID-19 analysis, evaluating them based on methodology, data, performance, and prerequisites for future research. By examining diverse data types and diagnostic modalities, this research contributes to scientific understanding and knowledge of the multimodal application of DL and its effectiveness in diagnosis. We have implemented and analysed 11 deep learning models using COVID-19 image, text, and speech (ie, cough) data. Our analysis revealed that the MobileNet model achieved the highest accuracy of 99.97% for COVID-19 image data and 93.73% for speech data (i.e., cough). However, the BiGRU model demonstrated superior performance in COVID-19 text classification with an accuracy of 99.89%. The broader implications of this research suggest potential benefits for other domains and disciplines that could leverage deep learning techniques for image, text, and speech analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。