arXiv:2603.19539cs.CLcs.AI2026-03被引 2

构建真实药品说明书问答基准,评估大模型在监管与临床推理上的表现

FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

  • 联合FDA专家构建多阶段标注流程,生成高质量问答数据
  • 发现现有模型在事实依据、长文本检索和安全拒答上存在明显短板
  • 适用于医药领域大模型的监管级评估,适合医疗AI研发人员使用

我们提出一个由专家精心构建的真实世界基准,用于评估基于文档的问答任务,其背景源于美国食品药品监督管理局(FDA)的仿制药评估。药品说明书包含丰富但异构的临床与监管信息,使当前语言模型难以准确回答问题。在与FDA监管评估人员合作的基础上,我们构建了FDARxBench,并设计多阶段流程生成高质量、专家标注的问答样本,涵盖事实型、多跳推理及拒绝回答任务,同时制定开卷与闭卷评估协议。对专有及开源模型的实验揭示了在事实锚定、长上下文检索和安全拒答行为方面的显著差距。尽管以FDA仿制药评估为出发点,该基准也为药品说明书理解的监管级评估提供了坚实基础,可用于评估大模型在药物标签问题上的表现。

原文摘要 · Abstract (English)

We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain rich but heterogeneous clinical and regulatory information, making accurate question answering difficult for current language models. In collaboration with FDA regulatory assessors, we introduce FDARxBench, and construct a multi-stage pipeline for generating high-quality, expert curated, QA examples spanning factual, multi-hop, and refusal tasks, and design evaluation protocols to assess both open-book and closed-book reasoning. Experiments across proprietary and open-weight models reveal substantial gaps in factual grounding, long-context retrieval, and safe refusal behavior. While motivated by FDA generic drug assessment needs, this benchmark also provides a substantial foundation for challenging regulatory-grade evaluation of label comprehension. The benchmark is designed to support evaluation of LLM behavior on drug-label questions.

药品说明问答系统大模型评估医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。