arXiv:2605.28422cs.CVcs.AI2026-05

医学多模态模型用双监督提升推理能力与可解释性

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

论文配图:VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
图 1 · 摘自论文原文
  • 用视觉与语义双重监督训练隐式推理过程
  • 在7个基准上超越现有方法,接近万亿参数闭源模型
  • 推理时零开销,事后可加解释模块

隐式推理通过连续隐状态进行推理,避免了链式思维的语言瓶颈和计算开销,适用于医学视觉问答。但现有方法存在模态崩溃、视觉监督不足及训练-推理不一致问题,且隐状态不可解释,难以用于临床。本文提出VITAL框架,通过视觉-语义双监督增强医学多模态大模型的隐空间推理:辅助文本解码器从隐状态重构推理链,视觉投影器从冻结的独立医学视觉编码器回归感兴趣区域特征。两者在推理时丢弃,无额外开销,但可事后重连实现双解释性,提供文本与视觉双重推理说明。构建了包含61,000样本、覆盖9种影像模态的数据集,规模超过以往医学视觉隐式推理数据集一个数量级。在7个基准上的实验表明,VITAL持续显著优于基线模型、所有隐式推理方法及使用更大规模数据训练的医学多模态模型,达到与万亿参数闭源模型相当的顶尖性能。

原文摘要 · Abstract (English)

Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead of chain-of-thought for medical VQA. However, existing methods suffer from modality collapse, insufficient visual supervision, and train-inference mismatch. Moreover, their opaque latent states offer no interpretability, which is critical in clinical applications. We propose VITAL, a latent-space reasoning framework for medical MLLMs with visual-semantic dual supervision: an auxiliary text decoder reconstructs reasoning chains from latent states, while a visual projector regresses ROI features from a frozen, independent medical vision encoder. Both modules are discarded at inference with zero overhead, yet can be re-attached post-hoc for dual interpretability, providing textual and visual explanations of the reasoning process without sacrificing efficiency. We construct a 61K dataset spanning 9 imaging modalities, exceeding prior medical visual latent reasoning datasets by an order of magnitude. Experiments on 7 benchmarks show that VITAL consistently and substantially outperforms the backbone, all latent reasoning baselines, and medical MLLMs trained on far larger data, achieving state-of-the-art results competitive with trillion-parameter proprietary models.

医学多模态隐式推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。