arXiv:2605.16215cs.AIcs.CL2026-05

首个可审计的临床大模型训练全流程,确保透明与可复现。

Fully Open Meditron: An Auditable Pipeline for Clinical LLMs

论文配图:Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
图 1 · 摘自论文原文
  • 构建从数据到模型的全链路开源管道,包含医生审核的数据集和训练框架。
  • 在医疗问答任务中,性能超越基线模型6.6个百分点,达到新SOTA。
  • 适合需要可验证、可审计的医疗AI研发团队使用。

临床决策支持系统(CDSS)需要可追溯、可审计的流程以实现严格、可复现的验证。然而当前基于大语言模型(LLM)的CDSS大多仍不透明。多数所谓‘开源’模型仅开放权重,却隐藏数据来源、清洗流程和生成路径。医学领域尚无真正完整的全开源(FO)模型。本文提出首个全开源临床大模型管道——Fully Open Meditron,涵盖医生审核的训练语料库、可复现的数据构建与训练框架,以及对齐应用的评估协议。该语料库整合8个公开医学问答数据集,并通过三类医生验证的合成数据扩展:考试式问答、基于46,469份临床指南的指南驱动问答、以及临床案例。管道实施全系统去污染、教师生成结果的黄金标签重采样,及四位医生组成的全程验证。评估采用大模型作为裁判的协议,在专家编写的临床案例上进行,与204名人类评审员校准。将该流程应用于五个基线模型(Apertus-70B/8B-Instruct、OLMo-2-32B-SFT、EuroLLM-22B/9B-Instruct),所有改进版本均优于其基线。其中,Apertus-70B-MeditronFO在综合医学基准上较其基线提升6.6分(47.2% → 53.8%),确立新的全开源最优表现。Gemma-3-27B-MeditronFO在58.6%的对比中胜过MedGemma,HealthBench得分58%优于基线55.9%。结果表明,全开源流程可在不牺牲可审计性与可复现性的前提下,实现顶尖领域性能。

原文摘要 · Abstract (English)

Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Most "open" models are open-weight only, releasing parameters while withholding the data provenance, curation procedures, and generation pipelines that determine model behavior. Fully Open (FO) models, which expose the complete training stack end-to-end, do not currently exist in medicine. We introduce Fully Open Meditron, the first fully open pipeline for building LLM-CDSS, comprising a clinician-audited training corpus, a reproducible data construction and training framework, and a use-aligned evaluation protocol. The corpus unifies eight public medical QA datasets into a normalized conversational format and expands coverage with three clinician-vetted synthetic extensions: exam-style QA, guideline-grounded QA derived from 46,469 clinical practice guidelines, and clinical vignettes. The pipeline enforces system-wide decontamination, gold-label resampling of teacher generations, and end-to-end validation by a four-physician panel. We evaluate using an LLM-as-a-judge protocol over expert-written clinical vignettes, calibrated against 204 human raters. We apply the recipe to five FO base models (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct). All MeditronFO variants are preferred over their bases. Apertus-70B-MeditronFO improves +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new FO SoTA. Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-as-a-judge comparisons and outperforms it on HealthBench (58% vs 55.9%). These results show that fully open pipelines can achieve state-of-the-art domain-specific performance without sacrificing auditability or reproducibility.

医疗AI可解释性大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。