首个公开的长篇医学问答评测基准,由医生标注真实患者问题答案。
A Benchmark for Long-Form Medical Question Answering
- 构建基于医生标注的真实患者医疗问题长回答评测集。
- 对比多种开源与闭源模型,发现开源模型表现接近顶尖闭源模型。
- 首次公开评估大模型在正确性、有用性、偏见等维度的表现数据。
当前缺乏针对大语言模型(LLMs)在长篇医学问答(QA)中表现的评测基准。现有医学QA评测多聚焦自动指标和选择题,难以反映真实临床场景的复杂性。此外,多数长篇回答评估研究为闭源,缺乏医生标注数据,导致结果不可复现。本文提出一个公开可用的新基准,包含真实消费者医疗问题及由医学专家标注的长回答。我们对多种开源与闭源医学及通用大模型进行了成对比较,评估其在正确性、有用性、有害性和偏见等方面的性能。同时开展大模型作为评判者(LLM-as-a-judge)的分析,探究人类判断与模型评估的一致性。初步结果显示,开源模型在医学问答中展现出与领先闭源模型相当的潜力。代码与数据:https://github.com/lavita-ai/medical-eval-sphere
原文摘要 · Abstract (English)
There is a lack of benchmarks for evaluating large language models (LLMs) in long-form medical question answering (QA). Most existing medical QA evaluation benchmarks focus on automatic metrics and multiple-choice questions. While valuable, these benchmarks fail to fully capture or assess the complexities of real-world clinical applications where LLMs are being deployed. Furthermore, existing studies on evaluating long-form answer generation in medical QA are primarily closed-source, lacking access to human medical expert annotations, which makes it difficult to reproduce results and enhance existing baselines. In this work, we introduce a new publicly available benchmark featuring real-world consumer medical questions with long-form answer evaluations annotated by medical doctors. We performed pairwise comparisons of responses from various open and closed-source medical and general-purpose LLMs based on criteria such as correctness, helpfulness, harmfulness, and bias. Additionally, we performed a comprehensive LLM-as-a-judge analysis to study the alignment between human judgments and LLMs. Our preliminary results highlight the strong potential of open LLMs in medical QA compared to leading closed models. Code & Data: https://github.com/lavita-ai/medical-eval-sphere
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。