arXiv:2604.08815cs.CV2026-04

让医学多模态模型基于多证据一致判断,提升诊断可信度。

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

  • 用影像统计、可解释性激活等信号强制多模态证据对齐
  • AUC提升至0.925,幻觉词减少78%,推理更简洁
  • 适合医疗AI可信决策场景,尤其关注模型可靠性

医学视觉语言模型在放射科任务中表现优异,但常因过度依赖单一模态而生成看似合理实则缺乏依据的结论。本文提出一种上下文对齐推理框架,在生成诊断结论前强制异质临床证据达成一致。该方法通过辐射组学统计、可解释性激活与词汇语义线索,为冻结的VLM注入结构化上下文信号。模型不再输出自由文本,而是生成包含支持证据、不确定性估计、局限性和安全提示的结构化输出。仅使用辅助信号效果有限;性能提升需通过上下文验证实现。在胸部X光数据集上的实验表明,上下文对齐使判别性能提升(AUC从0.918增至0.925),同时保持校准的不确定性。幻觉关键词数量从1.14降至0.25,推理解释长度由19.4词减至15.3词,模型置信度维持在0.68不变。跨数据集评估(CheXpert)显示模态信息量显著影响推理行为。结果表明,强制多证据一致可提升医学多模态推理的可靠性和可信度,且无需改变原模型架构。

原文摘要 · Abstract (English)

Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-reliance on a dominant modality. We introduce a context-aligned reasoning framework that enforces agreement across heterogeneous clinical evidence before generating diagnostic conclusions. The proposed approach augments a frozen VLM with structured contextual signals derived from radiomic statistics, explainability activations, and vocabulary-grounded semantic cues. Instead of producing free-form responses, the model generates structured outputs containing supporting evidence, uncertainty estimates, limitations, and safety notes. We observe that auxiliary signals alone provide limited benefit; performance gains emerge only when these signals are integrated through contextual verification. Experiments on chest X-ray datasets demonstrate that context alignment improves discriminative performance (AUC 0.918 to 0.925) while maintaining calibrated uncertainty. The framework also substantially reduces hallucinated keywords (1.14 to 0.25) and produces more concise reasoning explanations (19.4 to 15.3 words) without increasing model confidence (0.70 to 0.68). Cross-dataset evaluation on CheXpert further reveals that modality informativeness significantly influences reasoning behavior. These results suggest that enforcing multi-evidence agreement improves both reliability and trustworthiness in medical multimodal reasoning, while preserving the underlying model architecture.

医学多模态可信AI上下文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。