让医学大模型按临床思维诊断胃肠道病变,提升准确率。
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
- 构建分层临床认知数据集,让模型学习专家诊断逻辑。
- 通过反事实强化学习,消除图像背景干扰,聚焦真正病变特征。
- 适合医疗AI研发者和临床辅助诊断系统开发者参考。
多模态大语言模型在医学图像分析中展现出巨大潜力,但在胃肠镜检查应用中仍受两大限制:通用模型推理与标准化临床认知路径不匹配,以及视觉特征与诊断结果间缺乏因果关联。本文提出临床认知对齐(CogAlign)框架,首先通过构建分层临床认知数据集并采用监督微调(SFT),将专家从解剖定位、形态评估到微血管分析的层级诊断逻辑内化至模型;其次,基于理论分析指出标准监督微调会收敛至虚假背景相关性,进而提出基于反事实的强化学习策略,通过病变掩码生成反事实正常样本,并以临床认知为中心设计奖励函数优化模型,强制其基于因果病变特征进行诊断。大量实验表明,该方法在多个基准上达到最先进性能,显著提升复杂临床场景下的诊断准确率。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastrointestinal endoscopy is currently hindered by two critical limitations: the misalignment between general model reasoning and standardized clinical cognitive pathways, and the lack of causal association between visual features and diagnostic outcomes. In this paper, we propose a novel Clinical-Cognitive-Aligned (CogAlign) framework to address these challenges. First, we endow the model with rigorous clinical analytical capabilities by constructing the hierarchical clinical cognition dataset and employing Supervised Fine-Tuning (SFT). Unlike conventional approaches, this strategy internalizes the hierarchical diagnostic logic of experts, ranging from anatomical localization and morphological evaluation to microvascular analysis, directly into the model. Second, to eliminate visual bias, we provide a theoretical analysis demonstrating that standard supervised tuning inevitably converges to spurious background correlations. Guided by this insight, we propose a counterfactual-driven reinforcement learning strategy to enforce causal rectification. By generating counterfactual normal samples via lesion masking and optimizing through clinical-cognition-centric rewards, we constrain the model to strictly ground its diagnosis in causal lesion features. Extensive experiments demonstrate that our approach achieves State-of-the-Art (SoTA) performance across multiple benchmarks, significantly enhancing diagnostic accuracy in complex clinical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。