用医学证据自动生成精准评分标准,提升医疗大模型评估可靠性
Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation
- 检索权威医学内容生成原子级事实,结合用户意图合成细粒度评分标准
- 在HealthBench和LLMEval-Med上比GPT-4o高7.8%胜率,评分差异提升至8.658
- 可自动指导医疗回复优化,使响应质量提升9.2%,适合医疗AI研发者
大型语言模型(LLMs)在临床决策支持中应用日益广泛,但幻觉与不安全建议可能直接危及患者安全。现有评估方法难以捕捉细微临床错误:通用指标和基于通用标准的LLM评判效果有限,而专家制定的细粒度评分标准成本高、难扩展。本文提出一种检索增强型多智能体框架,自动生成实例相关的评价标准。该方法通过分解检索到的内容为原子事实,并融合用户交互约束,构建细粒度评估准则。在HealthBench和LLMEval-Med上,本框架实现临床意图对齐(CIA)得分分别为50.20%和31.90%,显著优于GPT-4o基准,在中英文医疗评测中均表现一致提升。在HealthBench的判别测试中,本框架胜率高出GPT-4o 7.8个百分点,平均分差由4.972增至8.658。消融实验显示,各组件在不同数据集贡献不同,其中交互意图建模对临床标准覆盖最稳定。此外,生成的评分标准可指导回复优化,使响应质量提升9.2%。结果表明,自动化、知识驱动的评分标准生成为医疗大模型的评估与改进提供了可扩展基础。代码已公开于https://anonymous.4open.science/r/Automated-Rubric-Generation-E716/。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed by generic metrics and LLM judges using general criteria, while expert-authored fine-grained rubrics are expensive and difficult to scale. In this paper, we propose a retrieval-augmented multi-agent framework for automatically generating instance-specific evaluation rubrics. Our approach grounds evaluation in authoritative medical evidence by decomposing retrieved content into atomic facts and synthesizing them with user interaction constraints to form fine-grained evaluation criteria. Evaluated on HealthBench and LLMEval-Med, our framework achieves Clinical Intent Alignment (CIA) scores of 50.20% and 31.90%, significantly outperforming the GPT-4o baseline and showing consistent improvements across English and Chinese medical benchmarks. In discriminative tests on HealthBench, our rubrics achieve a 7.8% point higher win rate than GPT-4o and increase the mean score difference from 4.972 to 8.658. Ablation studies further show that individual components contribute differently across datasets, with interaction-intent modeling providing the most consistent contribution to clinical-criterion coverage. Beyond evaluation, our rubrics guide response refinement, improving response quality by 9.2%. These results suggest that automated, knowledge-grounded rubric generation provides a scalable foundation for evaluating and improving medical LLMs. The code is available at https://anonymous.4open.science/r/Automated-Rubric-Generation-E716/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。