评测大模型当医学问答裁判的可靠性,发现需考虑生成模型影响。
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
- 用大模型评估法语医学开放问答语义等价性,关注生成器影响。
- 领域适配和大型通用模型最接近专家判断,小模型经微调后表现提升。
- 轻量级微调可降低对生成器敏感性,适合资源有限的医疗场景。
由于需要专家标注,医学开放问答(OEQA)的自动评估仍具挑战性。本文评估大语言模型(LLMs)在法语医学开放问答中作为语义等价性裁判的可行性,比较了闭源通用模型与生物医学领域适配模型。结果表明,基于LLM的评判受答案生成模型显著影响,不同生成器间一致性差异明显。领域适配模型和大型通用模型与专家标注对齐度最高。进一步发现,仅用少量数据通过监督微调(SFT)和组相对策略优化(GRPO)对小型模型进行轻量适配,即可显著提升性能并降低对生成器的敏感性。总体而言,研究强调了评估中需考虑生成器因素,并表明经过精心适配的小模型可在低资源医疗环境中支持可扩展评估。
原文摘要 · Abstract (English)
Automatic evaluation of medical open-ended question answering (OEQA) remains challenging due to the need for expert annotations. We evaluate whether large language models (LLMs) can act as judges of semantic equivalence in French medical OEQA, comparing closed-access, general-purpose, and biomedical domain-adapted models. Our results show that LLM-based judgments are strongly influenced by the model that generated the answer, with agreement varying substantially across generators. Domain-adapted and large general-purpose models achieve the highest alignment with expert annotations. We further show that lightweight adaptation of a compact model using supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO) substantially improves performance and reduces generator sensitivity, even with limited data. Overall, our findings highlight the need for generator-aware evaluation and suggest that carefully adapted small models can support scalable evaluation in low-resource medical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。