比较大模型长短上下文生成文献综述的优劣,发现需专家修正。
LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs
- 用短长上下文对比生成文献综述,评估信息覆盖与连贯性。
- 长上下文提升信息广度但更易重复、遗漏关键文献。
- 适合研究者参考初稿,但必须人工精修才能发表。
本研究评估了基于Semantic Scholar和Arxiv数据源,使用大语言模型(LLMs)在短上下文与长上下文设置下生成的文献综述质量,探究上下文窗口对AI生成文献综述的影响及AI在文献综述写作中的作用。共20篇AI生成的综述由两名研究人员从15个维度进行评估。结果表明,AI生成的综述需人工审核才能达到学术出版标准。随着上下文窗口增大,模型可整合更广泛信息并保持长文本连贯性,但也加剧了内容重复、关键工作遗漏以及描述性过强而缺乏综合分析的问题。研究显示,AI生成的综述可提供基础概览,但必须经领域专家批判性评估与修订。未来研究应探索在不同领域中结合其他大模型与微调模型,采用人机协同的混合方法,以克服当前局限。
原文摘要 · Abstract (English)
Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。