arXiv:2602.06570cs.CL2026-02被引 11

Baichuan-M3让AI像医生一样主动问诊、推理并防幻觉,提升医疗决策可靠性。

Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making

  • 模拟医生问诊流程,主动追问以消除信息模糊
  • 在多个医学评测中超越GPT-5.2,尤其在推理与安全方面
  • 适合临床辅助、医学教育及医疗AI研发人员使用

我们提出Baichuan-M3,一种面向医疗场景的增强型大语言模型,旨在将被动问答模式转变为主动、临床级的决策支持。针对现有系统在开放式问诊中的局限性,该模型采用专用训练流程,模拟医生系统的诊疗工作流。核心能力包括:(i) 主动获取信息以解决歧义;(ii) 长周期推理,整合零散证据形成连贯诊断;(iii) 自适应幻觉抑制,保障结果的真实性。实证评估显示,Baichuan-M3在新提出的HealthBench-Hallu和ScanBench,以及HealthBench上均达到当前最优表现,显著优于GPT-5.2在临床问诊、建议与安全性方面的表现。模型已公开发布于https://huggingface.co/collections/baichuan-inc/baichuan-m3。

原文摘要 · Abstract (English)

We introduce Baichuan-M3, a medical-enhanced large language model engineered to shift the paradigm from passive question-answering to active, clinical-grade decision support. Addressing the limitations of existing systems in open-ended consultations, Baichuan-M3 utilizes a specialized training pipeline to model the systematic workflow of a physician. Key capabilities include: (i) proactive information acquisition to resolve ambiguity; (ii) long-horizon reasoning that unifies scattered evidence into coherent diagnoses; and (iii) adaptive hallucination suppression to ensure factual reliability. Empirical evaluations demonstrate that Baichuan-M3 achieves state-of-the-art results on HealthBench, the newly introduced HealthBench-Hallu and ScanBench, significantly outperforming GPT-5.2 in clinical inquiry, advisory and safety. The models are publicly available at https://huggingface.co/collections/baichuan-inc/baichuan-m3.

医疗AI大模型临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。