arXiv:2509.14830cs.CVcs.AI2025-09ICCV被引 2

用人体骨骼影像和病历数据,实现可解释的骨健康分类。

ProtoMedX: Towards Explainable Multi-Modal Prototype Learning for Bone Health Classification

  • 基于原型学习的多模态模型,结合影像与病历信息。
  • 多模态下准确率达89.8%,优于现有方法。
  • 决策过程可视觉化解释,适合临床可信应用。

骨健康研究对骨质减少症和骨质疏松症的早期发现与治疗至关重要。临床诊断通常依赖双能X线吸收检测(DEXA扫描)和患者病史。当前人工智能在该领域的应用仍在探索中。大多数成功方法仅使用视觉数据(如DEXA/X光图像)进行深度学习,注重预测精度,而忽视可解释性,常依赖事后分析输入贡献。本文提出ProtoMedX,一种融合腰椎DEXA扫描与患者记录的多模态模型。其原型基础架构天生具备可解释性,这对医疗应用极为关键,尤其在即将实施的欧盟《人工智能法案》背景下,可明确分析模型决策,包括错误判断。ProtoMedX在骨健康分类上达到领先性能,同时提供临床医生可理解的可视化解释。基于4,160名真实NHS患者的数据库,该模型在纯视觉任务中达87.58%准确率,在多模态版本中达89.8%,均超越已有公开方法。

原文摘要 · Abstract (English)

Bone health studies are crucial in medical practice for the early detection and treatment of Osteopenia and Osteoporosis. Clinicians usually make a diagnosis based on densitometry (DEXA scans) and patient history. The applications of AI in this field are ongoing research. Most successful methods rely on deep learning models that use vision alone (DEXA/X-ray imagery) and focus on prediction accuracy, while explainability is often disregarded and left to post hoc assessments of input contributions. We propose ProtoMedX, a multi-modal (multimodal) model that uses both DEXA scans of the lumbar spine and patient records. ProtoMedX's prototype-based architecture is explainable by design, which is crucial for medical applications, especially in the context of the upcoming EU AI Act, as it allows explicit analysis of model decisions, including incorrect ones. ProtoMedX demonstrates state-of-the-art performance in bone health classification while also providing explanations that can be visually understood by clinicians. Using a dataset of 4,160 real NHS patients, the proposed ProtoMedX achieves 87.58% accuracy in vision-only tasks and 89.8% in its multi-modal variant, both surpassing existing published methods.

可解释AI多模态骨健康原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。