arXiv:2511.16139cs.AI2025-11被引 4

用几何约束提升医疗大模型评分能力,让评估更贴合临床实际。

Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints

  • 将医疗标准转为多维矩阵,通过几何投影约束对齐临床思维。
  • 在Healthbench上使基础模型性能提升45%(全集)和85%(难例)。
  • 适合医疗AI评估、大模型对齐与临床决策系统研发者参考。

将大语言模型(LLMs)融入医疗实践具有变革潜力,但其真实临床应用受限于三大问题:(1)静态评估基准与动态临床认知需求不匹配;(2)难以适应持续演进的多源医疗标准;(3)传统奖励模型无法体现复杂多维的医疗质量评价标准。为此,我们提出MR-RML(多维量规导向奖励模型学习)框架,结合GPRC(几何投影参考约束),将医疗标准结构化为多视角矩阵,指导数据生成与模型优化。方法包含三项创新:(1)在训练全流程嵌入领域特定指南的医疗标准体系;(2)独立的多维奖励模型分解评价维度,实现从规则或LLM打分到内部化奖励建模的跃迁;(3)通过几何投影参考约束,将临床认知逻辑转化为数学正则化项,使评分梯度与临床推理一致,支持合成数据训练。在权威医学基准Healthbench上的大量评估表明,该方法显著提升基线Qwen-32B模型表现,全集提升45%,难例集提升85%。其结果达开源模型最优水平,全集得分62.7,难例集44.7,超越多数闭源模型。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) into medical practice offers transformative potential, yet their real-world clinical applicability remains constrained by critical alignment issues: (1) a misalignment between static evaluation benchmarks and the dynamic cognitive demands of clinical practice, (2) challenges in adapting to continuously evolving, multi-source medical standards, and (3) the limited capacity of conventional reward models to reflect nuanced, multi-dimensional medical quality criteria. To overcome these limitations, we introduce MR-RML (Multidimensional Rubric-oriented Reward Model Learning) with GPRC (Geometric Projection Reference Constraints)-a novel alignment framework that structured medical standards into a multi-perspective matrix to guide both data generation and model optimization. Our approach introduces three key innovations: (1) a medical standard system that embeds domain-specific guidelines throughout the training pipeline; (2) an independent multi-dimensional reward model that decomposes evaluation criteria, transitioning from rule-based or LLM-based scoring to internalized reward modeling for better evaluation performance; and (3) geometric projection reference constraints that translate clinical cognitive logic into mathematical regularization, aligning scoring gradients with clinical reasoning and facilitating training with synthetically generated data. Extensive evaluations on the authoritative medical benchmark Healthbench demonstrate that our method significantly boosts the performance of the base Qwen-32B model, with improvements of 45% on the full subset and 85% on the hard subset. It achieves state-of-the-art results among open-source LLMs, scoring 62.7 (full) and 44.7 (hard), while also surpassing the majority of closed-source models.

医疗AI大模型对齐奖励模型临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。