通过专科适配与临床指南强化,提升医学多模态模型的复杂推理能力。
MMedExpert-R1: Strengthening Multimodal Medical Reasoning via Domain-Specific Adaptation and Clinical Guideline Reinforcement
- 构建专科化LoRA模块,实现多领域初始化
- 基于临床指南设计优势函数,匹配真实诊疗逻辑
- 融合多专家模型,支持跨专科可靠推理
医学视觉语言模型在感知任务中表现优异,但在真实场景所需的复杂临床推理方面仍存在不足。现有强化学习方法面临深层推理数据稀缺、冷启动限制多专科对齐、标准算法无法建模临床推理多样性等问题。本文提出MMedExpert-R1,通过领域特定适应与临床指南强化解决上述挑战。构建包含10,000个样本的高质量数据集MMedExpert,覆盖四大专科并标注逐步推理过程。采用领域特定适应(DSA)生成专科化LoRA模块以提供多样化初始化;设计基于指南的优势函数(GBA),显式建模不同临床推理视角,契合真实诊断策略。冲突感知能力集成将多个专业专家融合为统一智能体,确保强跨专科对齐。全面实验表明,7B模型在MedXpert-MM上达到27.50分,在OmniMedVQA上达83.03分,显著优于现有方法,为可靠多模态医学推理系统奠定坚实基础。
原文摘要 · Abstract (English)
Medical Vision-Language Models (MedVLMs) excel at perception tasks but struggle with complex clinical reasoning required in real-world scenarios. While reinforcement learning (RL) has been explored to enhance reasoning capabilities, existing approaches face critical mismatches: the scarcity of deep reasoning data, cold-start limits multi-specialty alignment, and standard RL algorithms fail to model clinical reasoning diversity. We propose MMedExpert-R1, a novel reasoning MedVLM that addresses these challenges through domain-specific adaptation and clinical guideline reinforcement. We construct MMedExpert, a high-quality dataset of 10K samples across four specialties with step-by-step reasoning traces. Our Domain-Specific Adaptation (DSA) creates specialty-specific LoRA modules to provide diverse initialization, while Guideline-Based Advantages (GBA) explicitly models different clinical reasoning perspectives to align with real-world diagnostic strategies. Conflict-Aware Capability Integration then merges these specialized experts into a unified agent, ensuring robust multi-specialty alignment. Comprehensive experiments demonstrate state-of-the-art performance, with our 7B model achieving 27.50 on MedXpert-MM and 83.03 on OmniMedVQA, establishing a robust foundation for reliable multimodal medical reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。