arXiv:2412.02149cs.CLcs.IR2024-12

用大模型生成有深度对比的文献综述,提升科研效率

Leveraging Large Language Models for Comparative Literature Summarization with Reflective Incremental Mechanisms

  • 分步提取关键信息,逐步构建对比摘要
  • 在1000篇论文数据集上优于GPT-4等基线模型
  • 适合需要快速比较多篇论文的研究者

本文提出ChatCite,一种利用大语言模型生成对比性文献综述的新方法。现有摘要模型虽能生成简洁内容,但缺乏深层对比洞察。ChatCite通过多步推理机制,从论文中提取关键要素,增量式构建对比摘要,并借助反思记忆过程优化输出。我们在自建数据集CompLit-LongContext(含1000篇研究论文及标注对比摘要)上评估该方法,实验结果表明,其在ROUGE和新提出的G-Score等自动指标上均优于GPT-4、BART、T5和CoT等基线模型。人工评估进一步证实,ChatCite生成的摘要更具连贯性、洞察力和流畅性。本方法显著推进了自动化文献综述生成技术,为研究人员高效对比与整合科学成果提供了有力工具。

原文摘要 · Abstract (English)

In this paper, we introduce ChatCite, a novel method leveraging large language models (LLMs) for generating comparative literature summaries. The ability to summarize research papers with a focus on key comparisons between studies is an essential task in academic research. Existing summarization models, while effective at generating concise summaries, fail to provide deep comparative insights. ChatCite addresses this limitation by incorporating a multi-step reasoning mechanism that extracts critical elements from papers, incrementally builds a comparative summary, and refines the output through a reflective memory process. We evaluate ChatCite on a custom dataset, CompLit-LongContext, consisting of 1000 research papers with annotated comparative summaries. Experimental results show that ChatCite outperforms several baseline methods, including GPT-4, BART, T5, and CoT, across various automatic evaluation metrics such as ROUGE and the newly proposed G-Score. Human evaluation further confirms that ChatCite generates more coherent, insightful, and fluent summaries compared to these baseline models. Our method provides a significant advancement in automatic literature review generation, offering researchers a powerful tool for efficiently comparing and synthesizing scientific research.

文献综述大模型对比生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。