arXiv:2512.19691cs.AIstat.AP2025-12

用医生监督重审医学评分数据集,发现27%标签错误,重算后模型性能提升13.5个百分点。

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight

  • 构建医生参与的可扩展审核流程,重审由LLM辅助生成的医疗评分标签。
  • 27%测试标签存在错误或不可计算,重算后与医生标注一致率达74%。
  • 适合关注医疗AI评估可信度、模型训练数据质量的研究者和临床开发者。

机器学习基准的参考标签越来越多依赖大模型辅助生成,但其可靠性尚未充分验证。我们审计了部分由大模型辅助生成标签的医学评分计算基准MedCalc-Bench,开发了一套可扩展的医生在环监督机制进行重新评估。至少27%的测试标签可能存在错误或无法计算。在50个实例的医生验证子集中,重算后的标签与医生真实标注的一致率为74%(95%置信区间:60%-84%),而原始标签仅为20%(95%置信区间:11%-33%)。使用原始标签评估前沿大模型会低估其准确率16-23个百分点。在受控强化学习实验中,基于重算标签训练的模型在医生标注实例上表现优于原始标签训练模型13.5个百分点(95%置信区间:10.6%-16.6%),且该优势延伸至相关医疗任务。大模型辅助基准若未经主动治理,可能将系统性误差传播至评估与后续训练环节。

原文摘要 · Abstract (English)

Reference labels for machine-learning benchmarks are increasingly synthesized with LLM assistance, but their reliability remains underexamined. We audit MedCalc-Bench, a clinical benchmark for medical score computation whose labels were partly derived with LLM assistance, and develop a scalable physician-in-the-loop stewardship pipeline to reassess them. At least 27% of test labels are likely erroneous or incomputable. On a 50-instance subset validated by physicians, our recomputed labels agree with physician ground truth 74% of the time (95% CI, 60-84%) versus 20% for the originals (95% CI, 11-33%). Using original labels to evaluate frontier LLMs underestimates accuracy by 16-23 percentage points. In a controlled reinforcement-learning experiment, a model trained on recomputed labels outperforms one trained on originals by 13.5 percentage points (95% CI, 10.6-16.6%) on physician-labeled instances, and this advantage extends to related medical tasks. LLM-assisted benchmarks can propagate systematic errors into both evaluation and post-training unless actively stewarded.

医疗AI基准评估大模型治理医生监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。