评测大模型生成文献引用、摘要和综述的能力,发现仍存幻觉且表现因学科而异。
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
- 构建多维度评估框架,测试大模型在三类文献写作任务的表现。
- 顶尖模型仍会生成虚假参考文献,且摘要与人类写作存在事实偏差。
- 不同学科下模型表现差异显著,适合需人工校验的学术写作场景。
大型语言模型(LLMs)有望自动化文献综述中的文献搜集、组织与总结等复杂流程。然而,目前尚不明确其在生成全面可靠文献综述方面的能力。本研究提出一个框架,用于自动评估LLMs在三个关键任务中的表现:参考文献生成、文献摘要和综述撰写。引入多维评价指标,评估生成参考文献的幻觉率,并衡量文献摘要与综述在语义覆盖度和事实一致性上与人工写作的差距。实验结果表明,即使最先进的模型仍会生成幻觉性参考文献,且不同模型在各学科领域的综述写作能力存在差异。这些发现凸显了提升LLMs在学术文献综述自动化中可靠性的重要性和必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) have emerged as a potential solution to automate the complex processes involved in writing literature reviews, such as literature collection, organization, and summarization. However, it is yet unclear how good LLMs are at automating comprehensive and reliable literature reviews. This study introduces a framework to automatically evaluate the performance of LLMs in three key tasks of literature writing: reference generation, literature summary, and literature review composition. We introduce multidimensional evaluation metrics that assess the hallucination rates in generated references and measure the semantic coverage and factual consistency of the literature summaries and compositions against human-written counterparts. The experimental results reveal that even the most advanced models still generate hallucinated references, despite recent progress. Moreover, we observe that the performance of different models varies across disciplines when it comes to writing literature reviews. These findings highlight the need for further research and development to improve the reliability of LLMs in automating academic literature reviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。