Crase用结构化图搜索提升学术检索,更可控且成本更低。
Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch

- 基于1.5跳引文网络扩展种子论文,只保留有支持证据的引用边。
- 在50万论文语料上,召回率比大模型代理高3倍,成本仅为三分之一。
- 适合需要可解释性与低成本的学术研究者使用。
我们提出Crase,一种结构受限且可检查的学术深搜替代方案。不同于开放式的搜索循环,Crase仅对搜索引擎查询一次以获取种子论文,随后沿其1.5跳引文邻域进行扩展,剔除无蕴含支持的引文边,并通过考虑时效性的随机游走对剩余论文排序。这使得候选集、每篇论文的保留理由及停止条件在推理前均明确固定。在包含50万论文的arXiv语料库上,针对LitSearch及其他基准测试,Crase在召回率@50上比基于专有模型的深度研究代理最高提升3倍,同时成本约为其三分之一。
原文摘要 · Abstract (English)
We present Crase, a bounded and inspectable alternative to deep research agents for scholarly search. Instead of an open-ended search loop, Crase queries a search engine once for seed papers, expands them along their 1.5-hop citation neighborhood, prunes citation edges whose claims lack entailment support, and ranks the remaining papers with a recency-aware random walk. This makes the candidate set, the reason each paper is kept, and the stopping condition explicit and fixed before inference. On LitSearch and one further benchmarks over a 500K-paper arXiv corpus, Crase outperforms deep research agents built on proprietary models by up to 3$\times$ recall@50 at roughly a third of the cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。