arXiv:2510.24980cs.CVcs.AI2025-10

用可解释的AI模型提升压疮分期准确率与临床可信度

FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning

  • 通过多模态大模型结合自我反思机制,迭代优化诊断推理
  • 在PIID数据集上达到85%准确率,比之前方法高4个百分点
  • 生成自然语言解释,适合临床部署与医生信任决策

压疮(PU)是严重且普遍的医疗问题。准确划分压疮阶段(I-IV期)对治疗至关重要,但因视觉差异细微和主观判断差异,临床评估存在较大变异性。以往基于卷积神经网络(CNN)和视觉变换器(ViT)的AI方法虽取得良好准确率,但可解释性不足。本文提出FT-ARM(微调的代理式反思多模态语言模型),一种基于LLaMA 3.2 90B微调的多模态大模型,具备代理式自我反思机制,模拟临床医生的诊断复核过程,通过推理视觉特征与文本编码的临床知识,持续优化预测结果。在公开的压疮图像数据集(PIID)上,FT-ARM实现了85%的分类准确率,较先前基于CNN的模型提升4%。不同于以往仅依赖离线评估的研究,FT-ARM专为实时推理设计并测试,贴近实际部署场景。同时,其输出具有临床依据的自然语言解释,显著提升可解释性与医生信任度。通过融合微调与跨模态反思推理,该模型提升了自动化伤口评估系统的可靠性、透明度与临床适用性,满足压疮分期一致性与可解释性的关键需求,助力患者护理改善。

原文摘要 · Abstract (English)

Pressure ulcers (PUs) are a serious and prevalent healthcare concern. Accurate classification of PU severity (Stages I-IV) is essential for proper treatment but remains challenging due to subtle visual distinctions and subjective interpretation, leading to variability among clinicians. Prior AI-based approaches using Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) achieved promising accuracy but offered limited interpretability. We present FT-ARM (Fine-Tuned Agentic Reflection Multimodal model), a fine-tuned multimodal large language model (MLLM) with an agentic self-reflection mechanism for pressure ulcer severity classification. Inspired by clinician-style diagnostic reassessment, FT-ARM iteratively refines its predictions by reasoning over visual features and encoded clinical knowledge from text, enhancing both accuracy and consistency. On the publicly available Pressure Injury Image Dataset (PIID), FT-ARM, fine-tuned from LLaMA 3.2 90B, achieved 85% accuracy in classifying PU stages I-IV, surpassing prior CNN-based models by +4%. Unlike earlier CNN/ViT studies that relied solely on offline evaluations, FT-ARM is designed and tested for live inference, reflecting real-time deployment conditions. Furthermore, it produces clinically grounded natural-language explanations, improving interpretability and trust. By integrating fine-tuning and reflective reasoning across multimodal inputs, FT-ARM advances the reliability, transparency, and clinical applicability of automated wound assessment systems, addressing the critical need for consistent and explainable PU staging to support improved patient care.

压疮识别多模态模型可解释AI医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。