评测大模型对英国公共卫生信息的掌握程度,发现顶尖模型在选择题中准确率超90%。
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
- 构建包含8000+问题的公共健康评测集PubHealthBench,自动从政府文件生成题目
- 最新闭源模型在选择题中准确率超90%,优于简单搜索的人类
- 自由回答任务表现差,无模型得分超过75%,需警惕生成内容风险
随着大语言模型(LLMs)日益普及,了解其在特定领域中的知识水平对实际应用至关重要,尤其在医学和公共卫生领域,错误信息可能严重影响英国居民。尽管已有诸多医疗领域评测基准,但对公共健康领域的模型知识仍缺乏系统评估。为此,本文提出新基准PubHealthBench,涵盖8000多个问题,用于评估模型在多项选择题回答(MCQA)和自由回答两种形式下的表现。数据源自687份当前英国政府指导文件,通过自动化流程生成题目。在该基准上评估24个模型发现,最新闭源模型(GPT-4.5、GPT-4.1 和 o1)在MCQA任务中准确率超过90%,优于仅用搜索引擎快速检索的普通人类。但在自由回答任务中,所有模型表现下降,最高得分未达75%。因此,尽管前沿模型在结构化问答中表现出色,作为公共健康信息源仍需额外保障或工具支持,尤其是在开放生成场景下。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become widely accessible, a detailed understanding of their knowledge within specific domains becomes necessary for successful real world use. This is particularly critical in the domains of medicine and public health, where failure to retrieve relevant, accurate, and current information could significantly impact UK residents. However, while there are a number of LLM benchmarks in the medical domain, currently little is known about LLM knowledge within the field of public health. To address this issue, this paper introduces a new benchmark, PubHealthBench, with over 8000 questions for evaluating LLMs' Multiple Choice Question Answering (MCQA) and free form responses to public health queries. To create PubHealthBench we extract free text from 687 current UK government guidance documents and implement an automated pipeline for generating MCQA samples. Assessing 24 LLMs on PubHealthBench we find the latest proprietary LLMs (GPT-4.5, GPT-4.1 and o1) have a high degree of knowledge, achieving >90% accuracy in the MCQA setup, and outperform humans with cursory search engine use. However, in the free form setup we see lower performance with no model scoring >75%. Therefore, while there are promising signs that state of the art (SOTA) LLMs are an increasingly accurate source of public health information, additional safeguards or tools may still be needed when providing free form responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。