arXiv:2512.21583cs.AI2025-12被引 2

用逻辑树约束视觉语言模型,提升医学多模态诊断可信度

A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning

  • 结合视觉语言对齐与逻辑规则化推理,分步拆解诊断任务
  • 在MedXpertQA上诊断准确率提升,推理过程更可解释
  • 适合需要高可信度医疗AI的临床场景和研究者

随着大语言模型(LLMs)和视觉语言模型(VLMs)在医学领域的快速发展,单纯整合临床文本与医学影像并不能保证可靠推理。现有多模态模型常产生幻觉或不一致的推理链条,限制了临床信任。本文提出一种基于LLaVA的诊断框架,融合视觉语言对齐与逻辑正则化推理。系统包含文本与图像输入编码器、跨模态对齐投影模块、将诊断任务分解为步骤的推理控制器,以及将分步前提组装为可验证结论的逻辑树生成器。在MedXpertQA及其他基准上的评估表明,该方法在多模态任务中提升了诊断准确率,同时生成更具可解释性的推理轨迹,在纯文本任务上也保持竞争力。结果表明,该工作为构建可信多模态医疗AI提供了可行路径。

原文摘要 · Abstract (English)

With the rapid growth of large language models (LLMs) and vision-language models (VLMs) in medicine, simply integrating clinical text and medical imaging does not guarantee reliable reasoning. Existing multimodal models often produce hallucinations or inconsistent chains of thought, limiting clinical trust. We propose a diagnostic framework built upon LLaVA that combines vision-language alignment with logic-regularized reasoning. The system includes an input encoder for text and images, a projection module for cross-modal alignment, a reasoning controller that decomposes diagnostic tasks into steps, and a logic tree generator that assembles stepwise premises into verifiable conclusions. Evaluations on MedXpertQA and other benchmarks show that our method improves diagnostic accuracy and yields more interpretable reasoning traces on multimodal tasks, while remaining competitive on text-only settings. These results suggest a promising step toward trustworthy multimodal medical AI.

多模态医疗AI逻辑推理视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。