arXiv:2505.02703cs.CV2025-05中稿 · IEEE TMI 2025被引 14

用因果模型消除医学图像问答中的偏见,提升答案准确性。

Structure Causal Models and LLMs Integration in Medical Visual Question Answering

  • 构建视觉与文本交互的因果图结构,显式建模问题对视觉特征的影响。
  • 通过多变量重采样前门调整,消除图像与问题间的虚假相关性,准确率显著提升。
  • 结合多种提示策略,增强模型对复杂医学数据的理解能力,适合临床辅助决策场景。

医学视觉问答(MedVQA)旨在根据医学图像回答医学问题。然而,医学数据的复杂性导致难以观测的混杂因素,图像与问题之间的偏差不可避免,严重影响医学意义答案的推断。本文提出一种针对MedVQA任务的因果推断框架,有效消除图像与问题间的相对混杂效应,保障问答的精确性。首次引入新型因果图结构,表征视觉与文本元素间的交互关系,显式捕捉不同问题如何影响视觉特征。优化过程中,利用互信息发现虚假相关性,并提出多变量重采样前门调整方法,旨在基于真实因果相关性对齐特征。此外,设计融合多种提示形式的提示策略,提升模型对复杂医学数据的理解与准确回答能力。在三个MedVQA数据集上的大量实验表明:1)本方法显著提升MedVQA准确率;2)在复杂医学数据下实现真正的因果关联。

原文摘要 · Abstract (English)

Medical Visual Question Answering (MedVQA) aims to answer medical questions according to medical images. However, the complexity of medical data leads to confounders that are difficult to observe, so bias between images and questions is inevitable. Such cross-modal bias makes it challenging to infer medically meaningful answers. In this work, we propose a causal inference framework for the MedVQA task, which effectively eliminates the relative confounding effect between the image and the question to ensure the precision of the question-answering (QA) session. We are the first to introduce a novel causal graph structure that represents the interaction between visual and textual elements, explicitly capturing how different questions influence visual features. During optimization, we apply the mutual information to discover spurious correlations and propose a multi-variable resampling front-door adjustment method to eliminate the relative confounding effect, which aims to align features based on their true causal relevance to the question-answering task. In addition, we also introduce a prompt strategy that combines multiple prompt forms to improve the model's ability to understand complex medical data and answer accurately. Extensive experiments on three MedVQA datasets demonstrate that 1) our method significantly improves the accuracy of MedVQA, and 2) our method achieves true causal correlations in the face of complex medical data.

医学问答因果推理视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。