揭示生成式AI如何系统性伪造学术引用,识别出重复出现的假参考文献。
How unique are hallucinated citations offered by generative Artificial Intelligence models?
- 分析假引用的生成模式,发现其基于真实作者、期刊等元素重组。
- 在10篇AI生成文章中,9.2%的引用为幻觉,含30%重复出现的虚构条目。
- 适合关注学术诚信与AI生成内容可信度的研究者阅读。
本文研究生成式AI如何产生并传播虚假学术引用,聚焦于被归于Ben Williamson和Nelli Piattoeva的虚构论文《Education Governance and Datafication》。通过谷歌学术与谷歌搜索定位137篇相关源文献,分析该虚构引用的结构、重复性及后续引用情况。结果表明,幻觉引用并非随机编造,而是对真实作者、期刊、年份和关键词的有规律重组,近30%的案例存在重复。对ChatGPT 5-mini的结构化提问显示,模型在无验证时会基于学习到的模式重构合理引用,而非记忆事实。此外,检查10篇关于数据化与学校治理的AI生成论文发现,尽管多数引用真实或部分准确,仍有9.2%为幻觉,其中包含最常见虚构引用的完全匹配。研究揭示了生成式AI持续存在的学术诚信风险,表明联网型AI仍无法彻底消除虚假引用。
原文摘要 · Abstract (English)
This paper investigates how generative AI produces and propagates hallucinated academic references, focusing on the recurring non-existent citation 'Education Governance and Datafication' attributed to Ben Williamson and Nelli Piattoeva. Drawing on 137 accessible source papers identified through Google Scholar and Google searches, the study analyses the structure, recurrence, and onward citation of this phantom reference. It shows that hallucinated citations are not random inventions but patterned recombinations of real authors, journals, dates, and keywords, with duplication occurring in nearly 30% of cases. The paper also reports a structured interrogation of ChatGPT 5-mini about how it generates citations and finds that, absent verification, the model reconstructs plausible references from learned patterns rather than factual recall. Finally, ten AI-generated essays on datafication and school governance were examined: while most references were genuine or partly accurate, 9.2% remained hallucinated, including an exact match to the most common phantom citation. The findings highlight ongoing risks to academic integrity and show that web-enabled AI still does not fully eliminate fabricated references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。