测试大模型在生成科研文献时的幻觉频率,发现不同模型差异显著。
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
- 构建双任务评估框架:乱序标题与混合标题,检测模型生成错误引用
- 15个主流模型中,幻觉率最高达42%,部分模型误造文献比例超30%
- 适合科研人员、AI伦理研究者关注模型生成可靠性问题
语言模型在信息生成与整合中的作用日益重要,其对科学知识的表征需高度准确。核心挑战是幻觉——生成看似合理但实际虚假的信息,包括虚构引文和不存在的研究论文。此类不准确在要求高事实正确性的领域(如学术与教育)中极具风险。本文提出ArxEval评估流程,利用ArXiv作为数据源,设计两个任务(Jumbled Titles与Mixed Titles),评估大模型在生成科研文献时产生幻觉的频率。评估涵盖15个广泛使用的语言模型,提供它们在处理科学文献时可靠性的对比分析。
原文摘要 · Abstract (English)
Language Models [LMs] are now playing an increasingly large role in information generation and synthesis; the representation of scientific knowledge in these systems needs to be highly accurate. A prime challenge is hallucination; that is, generating apparently plausible but actually false information, including invented citations and nonexistent research papers. This kind of inaccuracy is dangerous in all the domains that require high levels of factual correctness, such as academia and education. This work presents a pipeline for evaluating the frequency with which language models hallucinate in generating responses in the scientific literature. We propose ArxEval, an evaluation pipeline with two tasks using ArXiv as a repository: Jumbled Titles and Mixed Titles. Our evaluation includes fifteen widely used language models and provides comparative insights into their reliability in handling scientific literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。