构建处方级评测基准,精准评估大模型用药推荐能力
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation

- 基于患者动态病程设计多选题,每题包含具体药物-剂量-途径三元组
- 16个模型表现差异大,最高准确率仅46.10%,说明任务极具挑战性
- 适合临床决策系统开发者、医疗AI研究人员参考使用
住院用药推荐需随患者病情变化反复选择具体药物、剂量和给药途径。现有基准将此任务简化为基于粗粒度药品编码的入院级预测,输入为多热编码的诊断与操作码,无法反映真实处方中逐时点、信息丰富的特性。本文提出RxEval,一个处方级评测基准,通过多选题形式评估大模型的处方能力:每道题给出详细患者画像和时间序列临床轨迹,要求从真实处方中选出正确的药物-剂量-途径组合,并以基于推理链扰动生成的患者相关干扰项作为选项。RxEval包含1,547道题目,覆盖584名患者、18个诊断类别和969种独特药物。对16个大模型的评估显示,该基准具有高难度和强区分度:模型F1得分在45.18至77.10之间,最佳精确匹配率仅为46.10%。错误分析表明,即使前沿模型也可能忽略明确的患者信息,无法推导出正确临床结论。
原文摘要 · Abstract (English)
Inpatient medication recommendation requires clinicians to repeatedly select specific medications, doses, and routes as a patient's condition evolves. Existing benchmarks formulate this task as admission-level prediction over coarse drug codes with multi-hot diagnostic and procedure code inputs, failing to capture the per-timepoint, information-rich nature of real prescribing. We propose RxEval, a prescription-level benchmark that evaluates LLM prescribing capability by multiple-choice questions: each question presents a detailed patient profile and time-ordered clinical trajectory, requiring selection of specific medication-dose-route triples from real prescriptions and patient-specific distractors generated via reasoning-chain perturbation. RxEval comprises 1,547 questions spanning 584 patients, 18 diagnostic categories, and 969 unique medications. Evaluation of 16 LLMs shows that RxEval is both challenging and discriminative: F1 ranges from 45.18 to 77.10 across models, and the best Exact Match is only 46.10%. Error analysis reveals that even frontier models may overlook stated patient information and fail to derive clinical conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。