8款AI聊天机器人查文献参考文献,准确率不足三成,Grok和DeepSeek表现最好。
Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate
- 测试8款AI在生成学术参考文献上的表现,按五项标准评估。
- 仅26.5%参考文献完全正确,近四成是错误或虚构的。
- 适合学生和教师了解AI在学术写作中的可靠性风险。
本研究评估了八款生成式AI聊天机器人(ChatGPT、Claude、Copilot、DeepSeek、Gemini、Grok、Le Chat、Perplexity)在免费版本下,于大学语境中生成学术参考文献的表现。共评估400条参考文献,涵盖健康、工程、实验科学、社会科学和人文学科五大知识领域,采用标准化提示。每条参考文献依据作者、年份、标题、来源、位置、文献类型、出版年限及错误数五个关键维度评分。结果显示,仅有26.5%的参考文献完全正确,33.8%部分正确,39.8%存在错误或完全虚构。Grok与DeepSeek是唯一未生成虚假参考文献的模型,而Copilot、Perplexity与Claude的幻觉率最高。此外,各模型更倾向于生成图书类参考文献,但期刊文章的伪造率显著更高。多个模型间提供的来源高度重叠,尤其是DeepSeek、Grok、Gemini与ChatGPT之间。研究揭示当前AI模型存在结构性缺陷,警示学生盲目使用风险,强调高校需加强信息素养与批判性使用AI能力。
原文摘要 · Abstract (English)
This study analyzes the performance of eight generative artificial intelligence chatbots -- ChatGPT, Claude, Copilot, DeepSeek, Gemini, Grok, Le Chat, and Perplexity -- in their free versions, in the task of generating academic bibliographic references within the university context. A total of 400 references were evaluated across the five major areas of knowledge (Health, Engineering, Experimental Sciences, Social Sciences, and Humanities), based on a standardized prompt. Each reference was assessed according to five key components (authorship, year, title, source, and location), along with document type, publication age, and error count. The results show that only 26.5% of the references were fully correct, 33.8% partially correct, and 39.8% were either erroneous or entirely fabricated. Grok and DeepSeek stood out as the only chatbots that did not generate false references, while Copilot, Perplexity, and Claude exhibited the highest hallucination rates. Furthermore, the chatbots showed a greater tendency to generate book references over journal articles, although the latter had a significantly higher fabrication rate. A high degree of overlap was also detected among the sources provided by several models, particularly between DeepSeek, Grok, Gemini, and ChatGPT. These findings reveal structural limitations in current AI models, highlight the risks of uncritical use by students, and underscore the need to strengthen information and critical literacy regarding the use of AI tools in higher education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。