arXiv:2602.04982cs.CL2026-02被引 5

自动化评估医学问答与引用质量,提升大模型生成内容可信度。

BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations

  • 基于事实片段构建多维度评估框架,涵盖完整性和正确性。
  • 实验证明自动评估与人工评价高度相关,准确率超85%。
  • 适合医学AI研发者、文献审查人员使用,支持开源工具链。

随着大语言模型在医学问答中的广泛应用,评估其生成答案的质量及支撑性引文的可靠性变得至关重要。由于需要专家判断是否与科学文献一致,且医学术语复杂,文本生成质量评估仍是挑战。本文提出BioACE,一个自动化医学问答与引文评估框架。该框架从完整性、正确性、精确率和召回率等多个维度,对比生成答案与真实事实片段(ground-truth nuggets)进行评估。我们开发了自动化方法评估各项指标,并通过大量实验分析其与人工评估的相关性。此外,还采用自然语言推理(NLI)、预训练语言模型及大模型等技术,评估生成答案所引用的生物医学文献证据质量。通过详尽实验与分析,我们提供了最佳的医学问答与引文评估方案,作为BioACE(https://github.com/deepaknlp/BioACE)开源包的一部分。

原文摘要 · Abstract (English)

With the increasing use of large language models (LLMs) for generating answers to biomedical questions, it is crucial to evaluate the quality of the generated answers and the references provided to support the facts in the generated answers. Evaluation of text generated by LLMs remains a challenge for question answering, retrieval-augmented generation (RAG), summarization, and many other natural language processing tasks in the biomedical domain, due to the requirements of expert assessment to verify consistency with the scientific literature and complex medical terminology. In this work, we propose BioACE, an automated framework for evaluating biomedical answers and citations against the facts stated in the answers. The proposed BioACE framework considers multiple aspects, including completeness, correctness, precision, and recall, in relation to the ground-truth nuggets for answer evaluation. We developed automated approaches to evaluate each of the aforementioned aspects and performed extensive experiments to assess and analyze their correlation with human evaluations. In addition, we considered multiple existing approaches, such as natural language inference (NLI) and pre-trained language models and LLMs, to evaluate the quality of evidence provided to support the generated answers in the form of citations into biomedical literature. With the detailed experiments and analysis, we provide the best approaches for biomedical answer and citation evaluation as a part of BioACE (https://github.com/deepaknlp/BioACE) evaluation package.

医学AI评估框架引文验证大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。