arXiv:2510.22803cs.CV2025-10被引 1

让医学影像问答更可信:多组件可解释系统提升诊断透明度

MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering

  • 融合五项可解释技术,从提问重写到链式推理全程可视化
  • 在500张病理图像上综合得分达0.683,是基线的1.8倍
  • 每张图精准定位3-5个诊断区域,解释文本平均57词且术语专业

可解释性对医疗视觉问答(VQA)系统的临床应用至关重要,因医生需要透明的推理过程才能信任AI诊断。我们提出MedXplain-VQA,一个整合五个可解释AI组件的完整框架,实现可解释的医学图像分析。该框架采用微调的BLIP-2骨干网络,结合医学查询重写、增强型Grad-CAM注意力、精确区域提取以及基于多模态语言模型的结构化链式思考推理。为评估系统,我们引入专用于医疗领域的评估框架,以临床相关指标替代传统NLP指标,包括术语覆盖率、临床结构质量和注意力区域相关性。在500个PathVQA病理图像样本上的实验显示,改进系统综合得分为0.683,远超基线方法的0.378;同时保持高推理置信度(0.890)。系统每样本识别3-5个诊断相关区域,生成的结构化解释平均57词,使用恰当临床术语。消融实验表明,查询重写带来最大初始提升,而链式思考推理支持系统性诊断流程。这些发现证明MedXplain-VQA具备成为可靠可解释医学VQA系统的潜力。未来工作将聚焦于与医学专家验证及大规模临床数据集的应用,以确保临床就绪。

原文摘要 · Abstract (English)

Explainability is critical for the clinical adoption of medical visual question answering (VQA) systems, as physicians require transparent reasoning to trust AI-generated diagnoses. We present MedXplain-VQA, a comprehensive framework integrating five explainable AI components to deliver interpretable medical image analysis. The framework leverages a fine-tuned BLIP-2 backbone, medical query reformulation, enhanced Grad-CAM attention, precise region extraction, and structured chain-of-thought reasoning via multi-modal language models. To evaluate the system, we introduce a medical-domain-specific framework replacing traditional NLP metrics with clinically relevant assessments, including terminology coverage, clinical structure quality, and attention region relevance. Experiments on 500 PathVQA histopathology samples demonstrate substantial improvements, with the enhanced system achieving a composite score of 0.683 compared to 0.378 for baseline methods, while maintaining high reasoning confidence (0.890). Our system identifies 3-5 diagnostically relevant regions per sample and generates structured explanations averaging 57 words with appropriate clinical terminology. Ablation studies reveal that query reformulation provides the most significant initial improvement, while chain-of-thought reasoning enables systematic diagnostic processes. These findings underscore the potential of MedXplain-VQA as a robust, explainable medical VQA system. Future work will focus on validation with medical experts and large-scale clinical datasets to ensure clinical readiness.

医学AI可解释性视觉问答病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。