arXiv:2502.14275cs.CLcs.LG2025-02被引 3

用结构化单跳判断评估大模型医学知识真实性和可靠性

Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments

  • 构建基于UMLS的医学知识判断数据集,通过二分类测试单跳事实记忆
  • 大模型在罕见病上准确率低,且对错误答案过度自信
  • 引入检索增强生成可提升准确性,改善决策不确定性

大型语言模型(LLMs)在各类下游任务中被广泛应用,但其直接回忆和应用医学事实的能力仍缺乏深入研究。现有医学问答基准多侧重复杂推理或多跳推理,难以区分模型的事实知识与推理能力。鉴于医疗应用高风险性,评估模型事实正确性至关重要。为此,本文提出医学知识判断数据集(MKJ),源自统一医学语言系统(UMLS),一个标准化生物医学词汇与知识图谱的综合库。通过二分类框架,MKJ让模型评估简洁的一跳陈述有效性,从而直接衡量其知识保留能力。实验表明,大模型在准确回忆医学事实方面表现不佳,不同语义类型间差异显著,尤其在罕见疾病上表现薄弱;且模型存在严重校准问题,常对错误答案表现出过高信心。为缓解此问题,我们探索了检索增强生成方法,证实其能有效提升事实准确性并降低决策不确定性。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely adopted in various downstream task domains. However, their abilities to directly recall and apply factual medical knowledge remains under-explored. Most existing medical QA benchmarks assess complex reasoning or multi-hop inference, making it difficult to isolate LLMs' inherent medical knowledge from their reasoning capabilities. Given the high-stakes nature of medical applications, where incorrect information can have critical consequences, it is essential to evaluate the factuality of LLMs to retain medical knowledge. To address this challenge, we introduce the Medical Knowledge Judgment Dataset (MKJ), a dataset derived from the Unified Medical Language System (UMLS), a comprehensive repository of standardized biomedical vocabularies and knowledge graphs. Through a binary classification framework, MKJ evaluates LLMs' grasp of fundamental medical facts by having them assess the validity of concise, one-hop statements, enabling direct measurement of their knowledge retention capabilities. Our experiments reveal that LLMs have difficulty accurately recalling medical facts, with performances varying substantially across semantic types and showing notable weakness in uncommon medical conditions. Furthermore, LLMs show poor calibration, often being overconfident in incorrect answers. To mitigate these issues, we explore retrieval-augmented generation, demonstrating its effectiveness in improving factual accuracy and reducing uncertainty in medical decision-making.

医学AI大模型评估事实性检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。