arXiv:2601.18706cs.AIcs.LG2026-01被引 7

Health-SCORE让医疗大模型评估更高效,降低人工成本。

Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs

  • 用可扩展框架自动构建医疗领域评分标准
  • 评估效果接近人工制定标准,开发成本大幅降低
  • 支持强化学习奖励与提示工程,适合医疗AI研发者

评分标准对评估开放性大模型输出至关重要,尤其在医疗等安全敏感领域。然而,高质量、领域特定的评分标准通常需要大量专家时间和开发成本,难以规模化。本文提出 Health-SCORE,一种通用且可扩展的基于评分标准的训练与评估框架,显著降低评分标准开发成本,同时保持性能。我们证明,Health-SCORE不仅可用于独立评估,还可作为结构化奖励信号,指导具有安全意识监督的强化学习,并能直接嵌入提示中,通过上下文学习提升响应质量。在开放式医疗任务中,Health-SCORE的评估效果接近人工制定的标准,同时大幅降低开发投入,使基于评分标准的评估与训练更具可扩展性。

原文摘要 · Abstract (English)

Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human expertise time and development cost, making rubric-based evaluation and training difficult to scale. In this work, we introduce Health-SCORE, a generalizable and scalable rubric-based training and evaluation framework that substantially reduces rubric development costs without sacrificing performance. We show that Health-SCORE provides two practical benefits beyond standalone evaluation: it can be used as a structured reward signal to guide reinforcement learning with safety-aware supervision, and it can be incorporated directly into prompts to improve response quality through in-context learning. Across open-ended healthcare tasks, Health-SCORE achieves evaluation quality comparable to human-created rubrics while significantly lowering development effort, making rubric-based evaluation and training more scalable.

医疗大模型评分标准强化学习提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。