arXiv:2601.18356cs.LG2026-01被引 1

让医学视觉语言模型用因果推理跨模态思考,更可信。

Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning

  • 用因果图和临床案例检索增强推理,避开表面相关性
  • 在放射报告生成等任务中准确率提升,抗分布偏移能力更强
  • 适合医疗AI可信决策、临床辅助诊断场景

医学视觉语言模型(VLMs)在诊断报告生成和图像-文本对齐上表现优异,但其推理机制仍以相关性为主,依赖表层统计关联,无法捕捉临床决策核心的因果病理生理机制。这导致模型脆弱、易幻觉且受数据集偏差影响。检索增强生成(RAG)虽能部分缓解此问题,但依赖语义相似性,引入新的虚假关联。我们提出多模态因果检索增强生成框架,将因果推断与多模态检索结合,从外部来源检索临床相关实例和因果图,使模型推理基于反事实与干预证据,而非仅依赖相关性。该方法应用于放射报告生成、诊断预测和视觉问答任务,在事实准确性、对分布偏移的鲁棒性及可解释性方面均有提升。结果表明,因果检索为实现超越模式匹配的可信医学多模态推理提供了可扩展路径,适用于高风险临床场景。

原文摘要 · Abstract (English)

Medical vision-language models (VLMs) achieve strong performance in diagnostic reporting and image-text alignment, yet their underlying reasoning mechanisms remain fundamentally correlational, exhibiting reliance on superficial statistical associations that fail to capture the causal pathophysiological mechanisms central to clinical decision-making. This limitation makes them fragile, prone to hallucinations, and sensitive to dataset biases. Retrieval-augmented generation (RAG) offers a partial remedy by grounding predictions in external knowledge. However, conventional RAG depends on semantic similarity, introducing new spurious correlations. We propose Multimodal Causal Retrieval-Augmented Generation, a framework that integrates causal inference principles with multimodal retrieval. It retrieves clinically relevant exemplars and causal graphs from external sources, conditioning model reasoning on counterfactual and interventional evidence rather than correlations alone. Applied to radiology report generation, diagnosis prediction, and visual question answering, it improves factual accuracy, robustness to distribution shifts, and interpretability. Our results highlight causal retrieval as a scalable path toward medical VLMs that think beyond pattern matching, enabling trustworthy multimodal reasoning in high-stakes clinical settings.

医学AI因果推理多模态可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。