用LLM构建智能文献综述系统,无需复杂检索即可高效获取关键信息。
Patience is all you need! An agentic system for performing scientific literature review
- 基于关键词检索与LLM信息提炼,自动完成文献搜索与摘要生成
- 稀疏检索方法在生物领域问答任务中逼近顶尖表现
- 适合科研人员快速撰写文献综述,提升研究效率
大型语言模型(LLMs)在多个学科领域已用于支持问答任务。尽管其对基础问题已有良好表现,但在需要专业知识或语义复杂的场景中仍显不足。科学探究常涉及文献检索、从全文中提取关键信息,并分析不同发现间的支持或矛盾关系。相关信息往往隐藏于论文全文而非摘要中,且句子理解需依赖上下文。我们构建了一个基于LLM的系统,可自动执行文献搜索与信息提炼。在已发布的生物相关文献基准测试上评估了基于关键词的检索与信息抽取能力。结果表明,稀疏检索方法在无需密集检索及其配套基础设施的前提下,性能接近当前最优水平。同时验证了提升文献覆盖范围的有效策略,为生成全面的文献综述提供支持。
原文摘要 · Abstract (English)
Large language models (LLMs) have grown in their usage to provide support for question answering across numerous disciplines. The models on their own have already shown promise for answering basic questions, however fail quickly where expert domain knowledge is required or the question is nuanced. Scientific research often involves searching for relevant literature, distilling pertinent information from that literature and analysing how the findings support or contradict one another. The information is often encapsulated in the full text body of research articles, rather than just in the abstracts. Statements within these articles frequently require the wider article context to be fully understood. We have built an LLM-based system that performs such search and distillation of information encapsulated in scientific literature, and we evaluate our keyword based search and information distillation system against a set of biology related questions from previously released literature benchmarks. We demonstrate sparse retrieval methods exhibit results close to state of the art without the need for dense retrieval, with its associated infrastructure and complexity overhead. We also show how to increase the coverage of relevant documents for literature review generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。