用AI自动估算伤口大小,帮医生远程监控慢性伤口。
WoundAmbit: Bridging State-of-the-Art Semantic Segmentation and Real-World Wound Care
- 用统一标准测试多种深度学习模型,确保公平对比。
- 所有模型每秒处理至少1张图,且伤口区域激活明显。
- 医生评估显示模型分割效果好,适合临床部署。
慢性伤口影响大量人群,尤其老年人和糖尿病患者,常伴随行动不便与多重健康问题。通过手机拍摄图像实现自动化伤口监测,可减少线下就诊次数,实现远程跟踪伤口尺寸。语义分割是该流程的核心,但伤口分割在医学影像研究中仍不充分。为此,我们基准测试了通用视觉、医学影像及公开伤口挑战赛中的顶尖深度学习模型。为保证公平比较,我们统一训练流程、数据增强和评估方式,并采用交叉验证以降低划分偏差。同时评估实际部署因素,包括对分布外伤口数据集的泛化能力、计算效率与可解释性。此外,提出基于参考物的方法,将AI生成的掩码转换为临床相关的伤口尺寸估计,并基于医生评估,对五个最佳架构进行评价。总体而言,基于Transformer的TransNeXt展现出最强泛化能力。尽管推理时间有差异,所有模型在CPU上均实现每秒至少处理1张图像,满足应用需求。可解释性分析通常显示伤口区域激活显著,突出关注临床相关特征。专家评估表明所有模型掩码质量高,其中VWFormer与ConvNeXtS主干表现最优。尺寸估计精度在各模型间相近,预测结果与专家标注高度一致。最后,我们展示了所提出的AI驱动伤口尺寸估计框架WoundAmbit如何集成至定制化远程医疗系统中。
原文摘要 · Abstract (English)
Chronic wounds affect a large population, particularly the elderly and diabetic patients, who often exhibit limited mobility and co-existing health conditions. Automated wound monitoring via mobile image capture can reduce in-person physician visits by enabling remote tracking of wound size. Semantic segmentation is key to this process, yet wound segmentation remains underrepresented in medical imaging research. To address this, we benchmark state-of-the-art deep learning models from general-purpose vision, medical imaging, and top methods from public wound challenges. For a fair comparison, we standardize training, data augmentation, and evaluation, conducting cross-validation to minimize partitioning bias. We also assess real-world deployment aspects, including generalization to an out-of-distribution wound dataset, computational efficiency, and interpretability. Additionally, we propose a reference object-based approach to convert AI-generated masks into clinically relevant wound size estimates and evaluate this, along with mask quality, for the five best architectures based on physician assessments. Overall, the transformer-based TransNeXt showed the highest levels of generalizability. Despite variations in inference times, all models processed at least one image per second on the CPU, which is deemed adequate for the intended application. Interpretability analysis typically revealed prominent activations in wound regions, emphasizing focus on clinically relevant features. Expert evaluation showed high mask approval for all analyzed models, with VWFormer and ConvNeXtS backbone performing the best. Size retrieval accuracy was similar across models, and predictions closely matched expert annotations. Finally, we demonstrate how our AI-driven wound size estimation framework, WoundAmbit, is integrated into a custom telehealth system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。