SciRAG让论文检索更智能,能自动规划、引用文献并生成连贯结论。
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
- 动态切换串行与并行检索,适应复杂问题。
- 利用引用关系筛选文档,提升信息可信度。
- 按结构化提纲生成回答,确保逻辑清晰可追溯。
科学文献的快速增长对大规模知识整合系统提出了更高要求。尽管现有的检索增强生成(RAG)方法提升了科学信息获取效率,但仍存在忽视引用图结构、难以应对复杂查询、生成内容碎片化且难验证等问题。本文提出 SciRAG——一个开源的科学文献探索框架,通过三项创新解决上述问题:(1) 自适应检索,灵活切换串行与并行证据收集;(2) 引用感知的符号推理,利用引用图组织和过滤支持性文献;(3) 提纲引导的合成机制,规划、批判与优化答案以保证连贯性与透明溯源。在 QASA、ScholarQA 等多个基准上的实验表明,SciRAG 在事实准确率与合成质量上均优于现有系统,为可靠的大规模科学知识聚合奠定了新基础。
原文摘要 · Abstract (English)
The accelerating growth of scientific publications has intensified the need for scalable, trustworthy systems to synthesize knowledge across diverse literature. While recent retrieval-augmented generation (RAG) methods have improved access to scientific information, they often overlook citation graph structure, adapt poorly to complex queries, and yield fragmented, hard-to-verify syntheses. We introduce SciRAG, an open-source framework for scientific literature exploration that addresses these gaps through three key innovations: (1) adaptive retrieval that flexibly alternates between sequential and parallel evidence gathering; (2) citation-aware symbolic reasoning that leverages citation graphs to organize and filter supporting documents; and (3) outline-guided synthesis that plans, critiques, and refines answers to ensure coherence and transparent attribution. Extensive experiments across multiple benchmarks such as QASA and ScholarQA demonstrate that SciRAG outperforms prior systems in factual accuracy and synthesis quality, establishing a new foundation for reliable, large-scale scientific knowledge aggregation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。