arXiv:2502.06666cs.CLcs.AI2025-02NAACL被引 25

为医疗大模型评估设计多维度新框架,解决问答评价的盲区问题。

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

  • 提出多维度医疗LLM评估体系,融合开闭两种问答测试
  • 发现现有方法存在评估盲点与重复覆盖,影响结果可靠性
  • 发布新医疗基准CareQA并引入新指标缓解开集评估缺陷

当前大语言模型评估多依赖开放式或封闭式问答,前者捕捉话语能力但难判对错,后者评估事实性却缺乏表达力。二者常独立或联合使用,但其关联性尚未明了。本文聚焦医疗领域,构建全面的多维度评估体系,探索开闭式评测间的相关性与差异。研究揭示现有方法存在盲区与重叠。为此,我们发布新医疗基准CareQA,包含开闭两种变体,并提出新型开放评估指标“松散困惑度”(Relaxed Perplexity),以应对现有局限。

原文摘要 · Abstract (English)

Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evaluate the factuality of responses but lack expressiveness. Open-ended capture the model's capacity to produce discourse responses but are harder to assess for correctness. These two approaches are commonly used, either independently or together, though their relationship remains poorly understood. This work is focused on the healthcare domain, where both factuality and discourse matter greatly. It introduces a comprehensive, multi-axis suite for healthcare LLM evaluation, exploring correlations between open and close benchmarks and metrics. Findings include blind spots and overlaps in current methodologies. As an updated sanity check, we release a new medical benchmark --CareQA-- with both open and closed variants. Finally, we propose a novel metric for open-ended evaluations -- Relaxed Perplexity -- to mitigate the identified limitations.

医疗LLM评估框架困惑度基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。