arXiv:2507.09209cs.CV2025-07ICCV被引 4

用专家反馈动态修正医学视觉语言模型的不确定输出。

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

  • 通过不确定性检测识别不可靠回答,触发专家干预。
  • 4.2B模型在三基准上超越13B模型,仅需少量专家标注。
  • 无需重新训练,适合资源有限的临床部署场景。

视觉语言模型(VLMs)的快速发展推动了多模态医疗辅助系统的发展。然而,现有模型仍存在固有的概率不确定性,常产生错误或未经验证的回答,在医疗应用中后果严重。现有方法通过调整模型结构、使用高质量数据微调或偏好微调来提升医学视觉语言模型(MedVLM)性能,但这些依赖训练的策略成本高,且与临床专家意见对齐不足。为此,我们提出一种无需额外训练的专家闭环框架——专家控制的无分类器引导(Expert-CFG)。该框架引入不确定性估计策略,识别不可靠输出;随后检索相关参考文献,辅助专家标注意图关键词,并利用无分类器引导调整模型词元嵌入,确保输出正确且与专家标注一致。在三个医学视觉问答基准上的评估表明,该方法在仅使用4.2B参数和少量专家标注的情况下,优于参数量达13B的现有最先进模型,证明了其在资源受限环境下临床部署的可行性。

原文摘要 · Abstract (English)

The rapid advancements in Vision Language Models (VLMs) have prompted the development of multi-modal medical assistant systems. Despite this progress, current models still have inherent probabilistic uncertainties, often producing erroneous or unverified responses-an issue with serious implications in medical applications. Existing methods aim to enhance the performance of Medical Vision Language Model (MedVLM) by adjusting model structure, fine-tuning with high-quality data, or through preference fine-tuning. However, these training-dependent strategies are costly and still lack sufficient alignment with clinical expertise. To address these issues, we propose an expert-in-the-loop framework named Expert-Controlled Classifier-Free Guidance (Expert-CFG) to align MedVLM with clinical expertise without additional training. This framework introduces an uncertainty estimation strategy to identify unreliable outputs. It then retrieves relevant references to assist experts in highlighting key terms and applies classifier-free guidance to refine the token embeddings of MedVLM, ensuring that the adjusted outputs are correct and align with expert highlights. Evaluations across three medical visual question answering benchmarks demonstrate that the proposed Expert-CFG, with 4.2B parameters and limited expert annotations, outperforms state-of-the-art models with 13B parameters. The results demonstrate the feasibility of deploying such a system in resource-limited settings for clinical use.

医学AI视觉语言模型专家干预不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。