构建医疗多模态大模型,提升临床真实场景下的理解与推理能力
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
- 采用实体感知持续预训练,整合异构医学数据扩大知识覆盖
- 通过强化学习实现多步诊断推理,生成可验证的决策过程
- 融合用户偏好与证据驱动推理,降低幻觉并提升报告可靠性
我们提出 MedXIAOHE,一个面向真实临床应用的医疗视觉-语言基础模型,旨在推动通用医疗理解与推理。MedXIAOHE 在多个医学基准上达到当前最优表现,并超越多个领先闭源多模态系统。为实现此目标,我们设计了实体感知的持续预训练框架,有效组织异构医学语料,拓宽知识覆盖范围,减少长尾缺陷(如罕见病)。针对医疗专家级推理与交互,MedXIAOHE 通过强化学习引入多样化的医学推理模式,并结合工具增强的代理训练,实现可验证的多步诊断推理。为提升实际应用中的可靠性,模型整合用户偏好评分、基于证据的推理及低幻觉长篇报告生成,显著提高对医学指令的遵循度。本文发布该模型的设计选择、扩展洞见与评估框架,以期推动后续研究。
原文摘要 · Abstract (English)
We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world clinical applications. MedXIAOHE achieves state-of-the-art performance across diverse medical benchmarks and surpasses leading closed-source multimodal systems on multiple capabilities. To achieve this, we propose an entity-aware continual pretraining framework that organizes heterogeneous medical corpora to broaden knowledge coverage and reduce long-tail gaps (e.g., rare diseases). For medical expert-level reasoning and interaction, MedXIAOHE incorporates diverse medical reasoning patterns via reinforcement learning and tool-augmented agentic training, enabling multi-step diagnostic reasoning with verifiable decision traces. To improve reliability in real-world use, MedXIAOHE integrates user-preference rubrics, evidence-grounded reasoning, and low-hallucination long-form report generation, with improved adherence to medical instructions. We release this report to document our practical design choices, scaling insights, and evaluation framework, hoping to inspire further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。