arXiv:2603.01124cs.CVcs.AI2026-03被引 1

让医学视觉模型像医生一样看图推理,避免胡说八道。

ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models

  • 用临床假设引导图像区域生成推理链,让每步判断都有图可依。
  • 在三个医学问答和报告生成任务上,准确率显著高于现有方法。
  • 适合想提升医疗AI推理可信度的研究者和开发者。

医学视觉语言模型在临床决策支持中展现出巨大潜力,但常因缺乏对局部病理证据的充分依赖而产生事实性幻觉。现有医学对齐方法主要在输出层面通过偏好优化改进正确性,却使中间推理与视觉区域关联薄弱。尽管思维链(CoT)能增强多模态推理,但其仍以文本为中心,难以有效融合临床视觉线索。为此,我们提出ClinCoT,一种面向临床的视觉思维链框架,将偏好优化从输出层修正转变为视觉驱动的推理过程。我们构建了自动数据生成管道,通过假设驱动的区域提案进行推理,生成具有临床依据的偏好对。多个医学大模型评估器对每个回答进行排名并打分,这些排名作为监督信号训练目标模型。此外,引入基于评分的边界感知优化策略,结合偏好排序与分数差异,细化区域级推理轨迹。为保持模型策略演化过程中的对齐性,采用迭代学习机制动态重生成偏好数据。在三个医学视觉问答与报告生成基准上的实验表明,ClinCoT持续提升事实性定位能力,性能优于现有基于偏好的对齐方法。

原文摘要 · Abstract (English)

Medical Vision-Language Models have shown promising potential in clinical decision support, yet they remain prone to factual hallucinations due to insufficient grounding in localized pathological evidence. Existing medical alignment methods primarily operate at the response level through preference optimization, improving output correctness but leaving intermediate reasoning weakly connected to visual regions. Although chain-of-thought (CoT) enhances multimodal reasoning, it remains largely text-centric, limiting effective integration of clinical visual cues. To address this gap, we propose ClinCoT, a clinical-aware visual chain-of-thought framework that transforms preference optimization from response-level correction to visual-driven reasoning. We introduce an automatic data generation pipeline that constructs clinically grounded preference pairs through reasoning with hypotheses-driven region proposals. Multiple Med-LLMs evaluators rank and assign scores to each response, and these rankings serve as supervision to train the target model. We further introduce a scoring-based margin-aware optimization strategy that incorporates both preference ranking and score difference to refine region-level reasoning trajectories. To maintain alignment as the model's policy evolves during training, we adopt an iterative learning scheme that dynamically regenerates preference data. Extensive experiments on three medical VQA and report generation benchmarks demonstrate that ClinCoT consistently improves factual grounding and achieves superior performance compared with existing preference-based alignment methods.

医学AI视觉推理思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。