arXiv:2502.14714cs.AIcs.CL2025-02被引 7

检验ChatGPT生成医学知识的真假,发现它对药物基因准确率高,症状识别却较差。

From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT

  • 用生物医学本体框架验证LLM生成的疾病-药物/基因关联
  • 药物和基因识别准确率达88%-98%,症状识别仅49%-61%
  • 适合医疗AI安全评估、医学知识可信度研究者参考

大型语言模型的生成能力虽能加速科研任务,但其生成内容的真实性存疑。本文提出一种计算评估方法,通过构建疾病关联并利用生物医学本体的语义框架进行验证。以ChatGPT为模型,设计提示工程生成疾病与药物、症状、基因的关联,并在多个版本(如GPT-turbo、GPT-4)中测试。实验显示,疾病术语识别准确率为88%-97%,药物名称为90%-91%,基因信息为88%-98%;但症状术语识别率仅为49%-61%,因症状描述口语化、冗长,难以匹配本体正式语言。关联验证结果显示,疾病-药物与疾病-基因对的文献覆盖率分别为89%-91%,而症状相关关联仅49%-62%。

原文摘要 · Abstract (English)

The generative capabilities of LLM models offer opportunities for accelerating tasks but raise concerns about the authenticity of the knowledge they produce. To address these concerns, we present a computational approach that evaluates the factual accuracy of biomedical knowledge generated by an LLM. Our approach consists of two processes: generating disease-centric associations and verifying these associations using the semantic framework of biomedical ontologies. Using ChatGPT as the selected LLM, we designed prompt-engineering processes to establish linkages between diseases and related drugs, symptoms, and genes, and assessed consistency across multiple ChatGPT models (e.g., GPT-turbo, GPT-4, etc.). Experimental results demonstrate high accuracy in identifying disease terms (88%-97%), drug names (90%-91%), and genetic information (88%-98%). However, symptom term identification was notably lower (49%-61%), due to the informal and verbose nature of symptom descriptions, which hindered effective semantic matching with the formal language of specialized ontologies. Verification of associations reveals literature coverage rates of 89%-91% for disease-drug and disease-gene pairs, while symptom-related associations exhibit lower coverage (49%-62%).

医学生成大模型可信度知识验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。