测试5大AI平台医学文献引用检索准确率,发现近半数错误,平台差异显著。
Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study
- 用多指标综合评分法量化5个AI平台的文献引用检索错误
- 平均准确率仅0.29,47.8%的引用完全无法正确检索
- 不同AI平台和期刊间差异大,需人工核对引用数据
大型语言模型(LLMs)辅助医学文献检索可能导致错误引用,但此类错误尚未被严格量化。本研究定量评估了5种主流免费版LLM平台(Grok-2、ChatGPT GPT-4.1、Google Gemini Flash 2.5、Perplexity AI、DeepSeek GPT-4)在检索40篇随机选取的原始论文(每本期刊10篇)引用文献时的表现,这些论文发表于2024年1月至2025年7月的《英国医学杂志》(BMJ)、《美国医学会杂志》(JAMA)和《新英格兰医学杂志》(NEJM)。主要评估指标为多维度评分比(结合数字对象标识符有效性、PubMed ID、Google Scholar链接及相关性)和完全遗漏率(所有适用指标均失败的引用比例)。多变量回归分析显示,平台与评分比独立相关,期刊与完全遗漏率独立相关。5个平台平均评分比为0.29(标准差0.35,范围0-1.25),最高为Grok(0.57),最低为Gemini(0.11)。相较于BMJ,NEJM文章的评分比更低,完全遗漏率更高。结果表明LLMs整体表现有限,且在不同平台和期刊间存在显著差异。使用LLM辅助文献检索时,必须仔细核查参考文献数据。
原文摘要 · Abstract (English)
Large language models (LLMs) assisted literature retrieval may lead to erroneous references, but these errors have not been rigorously quantified. Therefore, we quantitatively assess errors in reference retrieval of widely used free-version LLM platforms and identify the factors associated with retrieval errors. We evaluated 2,000 references retrieved by 5 LLMs (Grok-2, ChatGPT GPT-4.1, Google Gemini Flash 2.5, Perplexity AI, and DeepSeek GPT-4) for 40 randomly-selected original articles (10 per journal) published Jan. 2024 to July 2025 from British Medical Journal (BMJ), Journal of the American Medical Association, and The New England Journal of Medicine (NEJM). Primary outcomes were a multimetric score ratio combining validity of digital object identifier, PubMed ID, Google-Scholar link, and relevance; and complete miss rate (proportion of references failing all applicable metrics). Multivariable regression was used to examine independent associations. LLM platforms completely failed to retrieve correct reference data 47.8% of the time. The average score ratio of the 5 LLM platforms was 0.29 (standard deviation, 0.35; range, 0-1.25), with a higher score ratio indicating a higher accuracy in retrieving relevant references and correct bibliographic data. The highest and lowest accuracies were achieved by Grok (0.57) and Genimi (0.11), respectively. Compared with BMJ, NEJM articles had lower score ratios and higher complete miss rates. Multivariable analysis shows LLM platforms and journals were independently associated with score ratios and complete miss rate, respectively. We show modest overall performance of LLMs and significant variability in retrieval accuracy across platforms and journals. LLM platforms and journals are associated with LLM's performance in retrieving medical literature. Bibliographic data should be carefully reviewed when using LLM-assisted literature retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。