提出新评估框架,精准检验生物医学大模型判断力
Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

- 用可审计的变异生成对比数据对,替代稀缺人工标注
- 从正确性、鲁棒性、格式合规三方面全面评估模型表现
- 发现微调+强化学习组合在医疗任务上最优
当高质量人工标注稀缺时,我们提出一种可扩展的、以有效性为导向的生物医学大模型判断器评估流程。首先,通过确定性的、基于指标的变异,增强现有生物医学基准数据集,生成可审计的偏好对。其次,不仅评估整体正确性,还从三个部署相关维度进行评估:与指标生成的黄金标准相比的正确性、重复随机采样下的鲁棒性,以及对指定输出格式的遵守情况。我们使用该流程评估 Llama-3.1-8B-Instruct 在四种模式下的表现:(1) 基线,直接使用指令模型;(2) SFT,仅使用基于蒸馏的监督微调;(3) RL,仅使用 GRPO 基于的强化学习;(4) SFT→RL,先进行 SFT 再进行 RL。基线和单阶段模式在结构化医学判别任务如 PICO 提取和临床计算中表现不佳,而 SFT→RL 模式在正确性、合规性和鲁棒性方面均表现最佳;优势集中于可分解任务(如 PICO、MedCalc),有时甚至达到或超越前沿模型水平。
原文摘要 · Abstract (English)
We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。