用多模型共识生成医学问答评分标准,提升回答临床相关性。
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

- 三模型独立生成评分条目,统一模型筛选语义一致项。
- 三态打分区分正确、遗漏与错误,错误得负分而非零分。
- 专家盲评证明其输出更符合临床需求,多基准排名第一。
基于可验证奖励的强化学习在数学和编程中表现优异,但开放式医学问题缺乏低成本验证方式:答案可能部分正确、不完整或含临床致命错误。由医生编写或验证的评分标准具备强临床依据,但每次调用专家成本过高。通过模型生成评分标准可实现监督扩展。本文提出ConRub-Med,确保评分反馈从构建到策略优化过程中保留有效区分度。针对每个提示,三个异构语言模型独立提出原子级标准;另一模型审查并仅保留三者均提供语义支持的标准。采用三态评分区分正确覆盖、信息缺失与错误陈述,错误项获得负分而非零分。在完整组相对策略优化(GRPO)中,若同一组内所有响应最终得分相同,则成对裁判仅当双方排序一致时才提供序列优势,不改变标量奖励;无平局情况则使用常规GRPO。在按问题匹配的盲测中,两名医学专家评估的完整流程输出比单一生成器产出的回应更具临床相关性。在所评估的开放模型中,ConRub-Med在九项基准中有六项排名第一,医疗与泛化平均分最高。利用包含5,166个提示的评分数据集,其在HealthBench-Hard上得分为$38.98 \pm 1.04$(均值±标准差),优于InfiMed-ORBIT的33.60(8,000样本)和37.30(28,000样本)。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers: responses may be partly correct, incomplete, or contain clinically consequential errors. Rubrics written or validated by physicians offer strong clinical grounding, but involving experts in every instance is costly. Model-generated rubrics make this supervision scalable. We introduce ConRub-Med to preserve useful distinctions as rubric feedback moves from construction to policy optimization. For each prompt, three heterogeneous language models propose atomic criteria independently; a separate model reviews them, retaining only criteria with semantic support from all three generators. Three-State scoring distinguishes correct coverage, missing information, and incorrect claims. Errors receive negative rather than zero credit. When every response in a complete Group Relative Policy Optimization (GRPO) group receives the same final reward, a pairwise judge provides sequence advantages only if both candidate orders agree, without changing the scalar rewards. Groups without ties use vanilla GRPO. In a blinded study matched by question, two medical experts rate panels from the full pipeline as more clinically relevant than panels produced by one generator. Across the evaluated open models, ConRub-Med ranks first on six of nine benchmarks and achieves the highest medical and generalization averages. Using the resulting rubric dataset of 5,166 prompts, it scores $38.98 \pm 1.04$ (mean $\pm$ SD) on HealthBench-Hard, compared with InfiMed-ORBIT's 33.60 with 8,000 samples and 37.30 with 28,000.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。