arXiv:2412.10982cs.AI2024-12被引 3

用知识图谱分析大模型的医学推理能力,发现通用模型更像人但不准,专业模型相反。

MedG-KRP: Medical Graph Knowledge Representation Probing

  • 构建医学知识图谱,可视化大模型如何连接医学概念
  • 人类评审显示GPT-4表现最佳,但与真实知识对比最差
  • 揭示通用模型与专业模型在医学推理中的差异,适合临床部署前评估

大型语言模型(LLMs)在医疗领域展现出强大潜力,其整合多源信息生成回答的能力类似于人类专家。然而,医学场景对推理准确性要求极高,现有基于多项选择题的评测基准常受质疑。为深入理解模型推理能力,本文提出一种基于知识图谱(KG)的方法,评估LLMs的生物医学推理能力。我们测试了GPT-4、Llama3-70b和专用医学模型PalmyraMed-70b,由60名医学生评审其生成的60个医学概念图,并与大型生物医学知识图谱BIOS对比。结果显示,人类评审中GPT-4表现最佳,但在与真实知识对比时误差最大;而专用模型PalmyraMed-70b则反之。本工作为可视化和验证大模型医学推理路径提供了有效工具,有助于安全、可靠地应用于临床。

原文摘要 · Abstract (English)

Large language models (LLMs) have recently emerged as powerful tools, finding many medical applications. LLMs' ability to coalesce vast amounts of information from many sources to generate a response-a process similar to that of a human expert-has led many to see potential in deploying LLMs for clinical use. However, medicine is a setting where accurate reasoning is paramount. Many researchers are questioning the effectiveness of multiple choice question answering (MCQA) benchmarks, frequently used to test LLMs. Researchers and clinicians alike must have complete confidence in LLMs' abilities for them to be deployed in a medical setting. To address this need for understanding, we introduce a knowledge graph (KG)-based method to evaluate the biomedical reasoning abilities of LLMs. Essentially, we map how LLMs link medical concepts in order to better understand how they reason. We test GPT-4, Llama3-70b, and PalmyraMed-70b, a specialized medical model. We enlist a panel of medical students to review a total of 60 LLM-generated graphs and compare these graphs to BIOS, a large biomedical KG. We observe GPT-4 to perform best in our human review but worst in our ground truth comparison; vice-versa with PalmyraMed, the medical model. Our work provides a means of visualizing the medical reasoning pathways of LLMs so they can be implemented in clinical settings safely and effectively.

医学AI知识图谱大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。