arXiv:2501.08208cs.CLcs.AI2025-01ACL被引 9

提出可自动化评估临床问答系统的三重指标,提升真实场景下模型可信度评测效率。

ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

  • 构建包含三指标的自动化评估体系:上下文相关性、拒绝准确性、对话忠实性
  • 在超200个白内障术后患者真实问题上验证,对话忠实性预测人类评分更准
  • 适配临床场景,助力医疗大模型持续迭代,开源数据与提示供研究复用

大语言模型在临床问答中展现巨大潜力,检索增强生成(RAG)是保障回答事实准确性的主流方法。然而现有自动化RAG评估指标在临床和对话场景表现不佳。人工评估成本高、难以扩展,不利于RAG系统持续迭代。为此,我们提出ASTRID——一种面向临床问答的自动化、可扩展三重评估框架,包含三个指标:上下文相关性(CR)、拒绝准确性(RA)和对话忠实性(CF)。其中CF为新设计,旨在更精准衡量模型响应与知识库的一致性,同时不惩罚自然对话特征。我们构建了超过200个真实患者问题的数据集,涵盖白内障手术术后随访场景,并补充急诊、临床及非临床域外问题。实证表明,CF在对话场景下对人类对忠实性的评价具有更强预测力。此外,结合CF、RA、CR的三重评估与临床医生判断高度一致,能有效识别不当、有害或无益响应。使用九种不同LLM的实验进一步证明,该三重指标可与人工评估高度吻合,具备集成至自动化评估流水线的潜力。我们公开所有提示与数据集,支持后续研究与发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive potential in clinical question answering (QA), with Retrieval Augmented Generation (RAG) emerging as a leading approach for ensuring the factual accuracy of model responses. However, current automated RAG metrics perform poorly in clinical and conversational use cases. Using clinical human evaluations of responses is expensive, unscalable, and not conducive to the continuous iterative development of RAG systems. To address these challenges, we introduce ASTRID - an Automated and Scalable TRIaD for evaluating clinical QA systems leveraging RAG - consisting of three metrics: Context Relevance (CR), Refusal Accuracy (RA), and Conversational Faithfulness (CF). Our novel evaluation metric, CF, is designed to better capture the faithfulness of a model's response to the knowledge base without penalising conversational elements. To validate our triad, we curate a dataset of over 200 real-world patient questions posed to an LLM-based QA agent during surgical follow-up for cataract surgery - the highest volume operation in the world - augmented with clinician-selected questions for emergency, clinical, and non-clinical out-of-domain scenarios. We demonstrate that CF can predict human ratings of faithfulness better than existing definitions for conversational use cases. Furthermore, we show that evaluation using our triad consisting of CF, RA, and CR exhibits alignment with clinician assessment for inappropriate, harmful, or unhelpful responses. Finally, using nine different LLMs, we demonstrate that the three metrics can closely agree with human evaluations, highlighting the potential of these metrics for use in LLM-driven automated evaluation pipelines. We also publish the prompts and datasets for these experiments, providing valuable resources for further research and development.

临床问答RAG评估大模型评测医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。