可解释模型Gyan在医学问答上达到87.1%准确率,优于主流大模型。
On the Performance of an Explainable Language Model on PubMedQA
- 采用解耦知识的组合式架构,提升可解释性与透明度。
- 在PubmedQA上达87.1%准确率,显著超越MedPrompt(GPT-4)与Med-PaLM 2。
- 无需大量训练资源,适合医疗等高可信场景部署。
大型语言模型(LLMs)在医学知识检索、推理和回答医学问题方面表现出色,接近医生水平。然而,这些模型缺乏可解释性,易产生幻觉,维护困难,且训练与推理需巨大算力。本文报告了基于新型架构的可解释语言模型Gyan在PubmedQA数据集上的结果。Gyan是一种组合式语言模型,其知识与模型解耦,具备可信赖、透明、不幻觉、无需大量训练和计算资源的特点,且易于跨领域迁移。Gyan-4.3在PubmedQA上取得87.1%的准确率,优于基于GPT-4的MedPrompt(82%)和Med-PaLM 2(81.8%)。未来将报告在MedQA、MedMCQA、MMLU-Medicine等数据集上的结果。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown significant abilities in retrieving medical knowledge, reasoning over it and answering medical questions comparably to physicians. However, these models are not interpretable, hallucinate, are difficult to maintain and require enormous compute resources for training and inference. In this paper, we report results from Gyan, an explainable language model based on an alternative architecture, on the PubmedQA data set. The Gyan LLM is a compositional language model and the model is decoupled from knowledge. Gyan is trustable, transparent, does not hallucinate and does not require significant training or compute resources. Gyan is easily transferable across domains. Gyan-4.3 achieves SOTA results on PubmedQA with 87.1% accuracy compared to 82% by MedPrompt based on GPT-4 and 81.8% by Med-PaLM 2 (Google and DeepMind). We will be reporting results for other medical data sets - MedQA, MedMCQA, MMLU - Medicine in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。