arXiv:2601.12632cs.CLcs.IR2026-01

新基准评估医学大模型事实性、抗干扰与偏见,更贴近真实临床场景。

BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models

  • 基于最新医学文献构建动态问答数据集,含2280对专家验证题
  • 药物说明书问答准确率最高达0.92,临床试验题仅0.36
  • 重点测试语言变体和潜在偏见,适合医疗AI评估者使用

大型语言模型(LLMs)在生物医学领域应用日益广泛,但现有基准数据集存在局限:多依赖静态或过时数据,难以反映生物医学知识的动态性与高风险性;且因与模型预训练语料重叠,存在数据泄露风险,同时忽视对语言变化鲁棒性和潜在人口统计偏见的评估。为此,我们提出BioPulse-QA,一个评估生物医学大模型在新发表文献(包括药物说明书、试验方案和临床指南)上表现的基准。该数据集包含2,280个专家验证的问答对及其扰动版本,涵盖抽取式与生成式两种形式。我们评估了GPT-4o、GPT-o1、Gemini-2.0-Flash和LLaMA-3.1 8B Instruct四款发布于基准文档出版日期前的模型。结果显示,GPT-o1在药物说明书上的宽松F1得分最高(0.92),其次为Gemini-2.0-Flash(0.90);临床试验是最大挑战,抽取式F1最低达0.36。性能差异在改写问题中大于拼写错误,偏见测试则显示差异不显著。BioPulse-QA提供了一种可扩展且临床相关的生物医学大模型评估框架。

原文摘要 · Abstract (English)

Objective: Large language models (LLMs) are increasingly applied in biomedical settings, and existing benchmark datasets have played an important role in supporting model development and evaluation. However, these benchmarks often have limitations. Many rely on static or outdated datasets that fail to capture the dynamic, context-rich, and high-stakes nature of biomedical knowledge. They also carry increasing risk of data leakage due to overlap with model pretraining corpora and often overlook critical dimensions such as robustness to linguistic variation and potential demographic biases. Materials and Methods: To address these gaps, we introduce BioPulse-QA, a benchmark that evaluates LLMs on answering questions from newly published biomedical documents including drug labels, trial protocols, and clinical guidelines. BioPulse-QA includes 2,280 expert-verified question answering (QA) pairs and perturbed variants, covering both extractive and abstractive formats. We evaluate four LLMs - GPT-4o, GPT-o1, Gemini-2.0-Flash, and LLaMA-3.1 8B Instruct - released prior to the publication dates of the benchmark documents. Results: GPT-o1 achieves the highest relaxed F1 score (0.92), followed by Gemini-2.0-Flash (0.90) on drug labels. Clinical trials are the most challenging source, with extractive F1 scores as low as 0.36. Discussion and Conclusion: Performance differences are larger for paraphrasing than for typographical errors, while bias testing shows negligible differences. BioPulse-QA provides a scalable and clinically relevant framework for evaluating biomedical LLMs.

医学AI大模型评测问答系统数据基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。