arXiv:2504.10452cs.CV2025-04被引 7

融合图像与位置信息,用Transformer提升伤口分类准确率

Integrating Vision and Location with Transformers: A Multimodal Deep Learning Framework for Medical Wound Analysis

  • 用Vision Transformer+小波变换提取图像特征,结合体表图定位
  • 优化算法使分类准确率达83.42%,比原始模型提升显著
  • 适合医疗图像分析、智能诊断系统开发者参考

准确识别急性和难愈性伤口是伤口诊断的关键步骤。高效分类模型可降低专科医生的时间和经济成本,并辅助制定最佳治疗方案。传统机器学习依赖特征选择,模型复杂且难以精准识别。近年来深度学习在伤口诊断中展现出潜力,但仍需提升效率与准确率。本研究提出一种基于深度学习的多模态分类框架,结合伤口图像与对应位置信息,对糖尿病、压疮、手术及静脉性溃疡等类型进行分类。构建体表图以提供位置数据,帮助医生更高效标注。模型采用Vision Transformer提取图像层次特征,引入离散小波变换(DWT)层捕捉高低频成分,再通过Transformer提取空间特征。神经元数量与权重向量优化采用三种群智能算法:魔物大猩猩优化器(MGTO)、改进灰狼优化器(IGWO)和狐狸优化算法。评估结果显示,使用优化算法可显著提升诊断准确率。仅用图像数据时,原体表图模型准确率为0.8123;结合图像与位置信息时为0.8007;使用优化模型后,准确率在0.7801至0.8342之间波动。

原文摘要 · Abstract (English)

Effective recognition of acute and difficult-to-heal wounds is a necessary step in wound diagnosis. An efficient classification model can help wound specialists classify wound types with less financial and time costs and also help in deciding on the optimal treatment method. Traditional machine learning models suffer from feature selection and are usually cumbersome models for accurate recognition. Recently, deep learning (DL) has emerged as a powerful tool in wound diagnosis. Although DL seems promising for wound type recognition, there is still a large scope for improving the efficiency and accuracy of the model. In this study, a DL-based multimodal classifier was developed using wound images and their corresponding locations to classify them into multiple classes, including diabetic, pressure, surgical, and venous ulcers. A body map was also created to provide location data, which can help wound specialists label wound locations more effectively. The model uses a Vision Transformer to extract hierarchical features from input images, a Discrete Wavelet Transform (DWT) layer to capture low and high frequency components, and a Transformer to extract spatial features. The number of neurons and weight vector optimization were performed using three swarm-based optimization techniques (Monster Gorilla Toner (MGTO), Improved Gray Wolf Optimization (IGWO), and Fox Optimization Algorithm). The evaluation results show that weight vector optimization using optimization algorithms can increase diagnostic accuracy and make it a very effective approach for wound detection. In the classification using the original body map, the proposed model was able to achieve an accuracy of 0.8123 using image data and an accuracy of 0.8007 using a combination of image data and wound location. Also, the accuracy of the model in combination with the optimization models varied from 0.7801 to 0.8342.

医学图像多模态Transformer伤口分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。