arXiv:2604.08326cs.AI2026-04ACL被引 2

用细粒度临床标准提升医疗大模型安全与准确率

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection

  • 通过医生制定的评分规则构建医疗指令数据集,实现标准显式注入
  • 训练多维度奖励模型,使模型在强化学习中精准区分安全与性能
  • 在真实医学评测中表现媲美闭源顶尖模型,适合医疗AI研发者使用

将大语言模型与高风险医疗标准对齐仍是重大挑战,主要源于粗粒度偏好信号与临床指南复杂多维特性之间的不一致。为此,我们提出ProMedical,一个基于细粒度临床标准的统一对齐框架。首先,通过人机协作流程构建ProMedical-Preference-50k数据集,为医疗指令添加严格、由医师制定的评分标准。基于该语料,我们提出显式标准注入范式,训练多维度奖励模型。不同于传统标量奖励模型,该方法显式解耦安全约束与通用能力,实现在强化学习中的精确引导。为严谨验证,我们建立基于双盲专家仲裁的ProMedical-Bench评测集。实证表明,采用ProMedical-RM指导的GRPO优化Qwen3-8B基础模型,整体准确率提升22.3%,安全合规性提高21.7%,性能接近专有前沿模型。此外,对齐策略在外部基准上具有强泛化能力,在UltraMedical上表现可比肩最先进模型。我们公开发布数据集、奖励模型与评测集,推动可复现的安全医疗对齐研究。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with high-stakes medical standards remains a significant challenge, primarily due to the dissonance between coarse-grained preference signals and the complex, multi-dimensional nature of clinical protocols. To bridge this gap, we introduce ProMedical, a unified alignment framework grounded in fine-grained clinical criteria. We first construct ProMedical-Preference-50k, a dataset generated via a human-in-the-loop pipeline that augments medical instructions with rigorous, physician-derived rubrics. Leveraging this corpus, we propose the Explicit Criteria Injection paradigm to train a multi-dimensional reward model. Unlike traditional scalar reward models, our approach explicitly disentangles safety constraints from general proficiency, enabling precise guidance during reinforcement learning. To rigorously validate this framework, we establish ProMedical-Bench, a held-out evaluation suite anchored by double-blind expert adjudication. Empirical evaluations demonstrate that optimizing the Qwen3-8B base model via ProMedical-RM-guided GRPO yields substantial gains, improving overall accuracy by 22.3% and safety compliance by 21.7%, effectively rivaling proprietary frontier models. Furthermore, the aligned policy generalizes robustly to external benchmarks, demonstrating performance comparable to state-of-the-art models on UltraMedical. We publicly release our datasets, reward models, and benchmarks to facilitate reproducible research in safety-aware medical alignment.

医疗大模型对齐方法奖励建模安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。