测试大模型写论文时引用的arXiv链接准确性,发现多数模型会编造不存在的文献。
ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
- 构建arXivBench基准,评估大模型在8个学科、5个计算机子领域生成准确论文链接的能力。
- Claude-3.5-Sonnet表现最佳,但多数模型在人工智能外领域错误率超40%。
- 适合关注学术写作可信度的研究者和期刊编辑参考。
大型语言模型(LLMs)在推理和问答任务中表现出色,但其生成内容常含事实性错误,仍是关键挑战。本研究评估了专有及开源大模型在生成包含准确arXiv链接的相关论文时的表现。评估结果揭示了严重的学术风险:大模型频繁生成错误的arXiv链接或引用不存在的论文,从根本上削弱了其对真实作者贡献的正确归因能力。为此,我们提出arXivBench,一个专门针对arXiv八个主要学科类别及计算机科学五个子领域的基准测试工具。结果显示,不同学科间准确性差异显著,Claude-3.5-Sonnet在生成相关且准确回答方面具有明显优势;值得注意的是,大多数模型在人工智能领域的表现远优于其他子领域。该基准为评估大模型在科研场景中的可靠性提供了标准化工具,有助于推动其在研究环境中更可信的应用。代码与数据集已公开于https://github.com/liningresearch/arXivBench 和 https://huggingface.co/datasets/arXivBenchLLM/arXivBench。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate strong capabilities in reasoning and question answering, yet their tendency to generate factually incorrect content remains a critical challenge. This study evaluates proprietary and open-source LLMs on generating relevant research papers with accurate arXiv links. Our evaluation reveals critical academic risks: LLMs frequently generate incorrect arXiv links or references to non-existent papers, fundamentally undermining their ability to properly attribute research contributions to the actual authors. We introduce arXivBench, a benchmark specifically designed to assess LLM performance across eight major subject categories on arXiv and five subfields within computer science, one of the most popular categories among them. Our findings show concerning accuracy variations across subjects, with Claude-3.5-Sonnet exhibiting a substantial advantage in generating both relevant and accurate responses. Notably, most LLMs perform significantly better in Artificial Intelligence than other subfields. This benchmark provides a standardized tool for evaluating LLM reliability in scientific contexts, promoting more dependable academic use in research environments. Our code and dataset are available at https://github.com/liningresearch/arXivBench and https://huggingface.co/datasets/arXivBenchLLM/arXivBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。