arXiv:2606.22151cs.IR2026-06

让论文检索会比较:自动分析研究差异与空白,帮新手快速摸清领域脉络。

Novelty-Aware Agentic Retrieval: Comparing Research Contributions Through Structured Multi-Step Reasoning

论文配图:Novelty-Aware Agentic Retrieval: Comparing Research Contributions Through Structured Multi-Step Reasoning
图 1 · 摘自论文原文
  • 用六步智能流程,把论文检索变成多步推理,挖掘研究间的异同点。
  • 在100篇论文上实现5项结构化对比能力,比传统检索高出3倍以上效果。
  • 适合刚进新领域的研究者,或需要系统性文献综述的项目团队。

科学文献检索不仅是找相关论文,更需理解论文间的关系、重叠、差异及方法-问题组合的缺失。标准检索增强生成(RAG)独立总结文档,丢失了这种比较信号。我们提出新型研究代理(Novelty-Aware Research Agent),基于六组件结构化多步推理框架,在RAG流程上叠加比较能力:查询分析、类ReAct检索循环、相关性排序、模式引导贡献提取、三轮对比代理和答案生成。系统不仅能返回相关论文,还能生成结构化对比结果:每篇论文的贡献记录、论文级重叠关系,以及问题×方法缺口矩阵。在100篇论文语料上,该系统实现了标准RAG无法支持的五项结构化对比功能;在三个主查询中,无论文出现在所有前三名集合中(平均两两交集Jaccard为0.12),十查询扩展评估中平均Jaccard为0.115,29篇被召回论文中有18篇仅对应单一查询。在作者标注的相关性评分下,排名器在主查询上达到Precision@5 1.000,nDCG@5 0.752,优于BM25、稠密和混合检索;十查询下,Precision@5为0.980(未饱和),nDCG@5为0.739。模式合规率在主查询为86.7%,十查询集为84.0%;20个采样空缺矩阵单元验证显示缺口精度为0.600。我们讨论了代理检索中的延迟-结构权衡,并指出语料规模、人工标注标签和有限独立评估是主要局限。

原文摘要 · Abstract (English)

Scientific literature search is an information retrieval (IR) task in which ranked lists are insufficient: a researcher entering a new area needs to know not only which papers are relevant, but how they relate, where they overlap, how they differ, and what problem-method combinations are absent. Standard retrieval-augmented generation (RAG) summarizes documents independently, discarding this comparative signal. We present the Novelty-Aware Research Agent, a prototype agentic retrieval system that layers structured multi-step reasoning on a RAG pipeline through six typed-contract components: query analysis, a ReAct-style retrieval loop, relevance ranking, schema-guided contribution extraction, a three-pass comparison agent, and answer generation. Beyond returning relevant papers, it produces structured comparison artifacts: per-paper contribution records, paper-level overlaps, and a problem x method gap matrix. On a 100-paper corpus, the system supports five structured comparison capabilities that a standard RAG baseline supports none of, while remaining query-sensitive: across three main queries no paper appears in all three top-5 sets (mean pairwise Jaccard 0.12), and an extended seven-query evaluation holds the pattern across ten queries (mean Jaccard 0.115, 18 of 29 retrieved papers query-exclusive). Under author-assigned graded relevance the ranker attains mean Precision@5 1.000 and nDCG@5 0.752 on the main queries, ahead of BM25, dense, and hybrid retrieval; over ten queries Precision@5 is non-saturated at 0.980 with nDCG@5 0.739. Schema compliance is 86.7% on the main queries and 84.0% over the ten-query set, and validating 20 sampled empty gap-matrix cells yields a gap precision of 0.600. We discuss the latency-structure trade-off in agentic retrieval and identify corpus scale, author-assigned labels, and limited independent evaluation as the main limitations.

文献检索智能代理研究洞察结构化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。