arXiv:2508.19319eess.IVcs.AI2025-08

融合视觉分析与医学知识检索,提升肌少症超声诊断准确率

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

  • 分层视觉理解+文本引导的特征融合,精准捕捉肌肉结构
  • 在自建数据集上达99%准确率,优于现有方法超10%
  • 适合临床辅助诊断研究者与多模态AI医疗开发者

基于超声的肌少症精准诊断仍面临影像线索微弱、标注数据有限及模型缺乏临床背景等挑战。本文提出MedVQA-TREE,一种多模态框架,包含分层图像解析模块、门控特征融合机制和新型多跳多查询检索策略。视觉模块通过解剖分类、区域分割与图式空间推理,捕捉粗粒度、中层及细粒度结构;门控融合机制选择性整合视觉特征与文本查询,临床知识则通过UMLS引导的管道访问PubMed及肌少症专用外部知识库。该模型在两个公开MedVQA数据集(VQA-RAD和PathVQA)及自建肌少症超声数据集上训练评估,最高诊断准确率达99%,性能超越先前最先进方法超过10%。结果表明,结合结构化视觉理解与引导式知识检索,可显著提升肌少症智能辅助诊断效果。

原文摘要 · Abstract (English)

Accurate sarcopenia diagnosis via ultrasound remains challenging due to subtle imaging cues, limited labeled data, and the absence of clinical context in most models. We propose MedVQA-TREE, a multimodal framework that integrates a hierarchical image interpretation module, a gated feature-level fusion mechanism, and a novel multi-hop, multi-query retrieval strategy. The vision module includes anatomical classification, region segmentation, and graph-based spatial reasoning to capture coarse, mid-level, and fine-grained structures. A gated fusion mechanism selectively integrates visual features with textual queries, while clinical knowledge is retrieved through a UMLS-guided pipeline accessing PubMed and a sarcopenia-specific external knowledge base. MedVQA-TREE was trained and evaluated on two public MedVQA datasets (VQA-RAD and PathVQA) and a custom sarcopenia ultrasound dataset. The model achieved up to 99% diagnostic accuracy and outperformed previous state-of-the-art methods by over 10%. These results underscore the benefit of combining structured visual understanding with guided knowledge retrieval for effective AI-assisted diagnosis in sarcopenia.

肌少症多模态医学问答知识检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。