arXiv:2508.08742cs.CLcs.AI2025-08被引 5

首个专评科学检索重排器的基准,助力提升大模型科研问答准确性。

SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs

  • 构建五大学科领域的重排器评测基准,设计三类挑战性样本。
  • 13个重排器在噪声、语义混淆、反事实场景中表现差异显著。
  • 适合研究科学问答、RAG系统优化与模型可解释性的学者使用。

科学文献问答是推动新科学发现的关键步骤。近年来,两阶段检索增强生成大语言模型(RAG-LLMs)在此领域取得显著进展。该框架中第二阶段的重排器尤为关键,因科学领域术语细微差异可能严重影响事实导向或知识密集型回答的准确性。尽管已有进步,但相关方法的潜力与局限仍不明确。本文提出科学重排器评测基准(SciRerankBench),用于评估RAG-LLMs系统中的重排器性能,覆盖五个科学领域。为严格测试重排器在抗噪声、语义歧义区分和事实一致性方面的表现,我们构建三类问题-上下文-答案对:含噪声上下文(NC)、语义相似但逻辑无关上下文(SSLI)、反事实上下文(CC)。通过对13个常用重排器在五类主流LLM上的系统评估,揭示其相对优劣。据我们所知,SciRerankBench是首个专门针对RAG-LLMs中重排器的评测基准,为未来研发提供宝贵洞见与指导。

原文摘要 · Abstract (English)

Scientific literature question answering is a pivotal step towards new scientific discoveries. Recently, \textit{two-stage} retrieval-augmented generated large language models (RAG-LLMs) have shown impressive advancements in this domain. Such a two-stage framework, especially the second stage (reranker), is particularly essential in the scientific domain, where subtle differences in terminology may have a greatly negative impact on the final factual-oriented or knowledge-intensive answers. Despite this significant progress, the potential and limitations of these works remain unexplored. In this work, we present a Scientific Rerank-oriented RAG Benchmark (SciRerankBench), for evaluating rerankers within RAG-LLMs systems, spanning five scientific subjects. To rigorously assess the reranker performance in terms of noise resilience, relevance disambiguation, and factual consistency, we develop three types of question-context-answer (Q-C-A) pairs, i.e., Noisy Contexts (NC), Semantically Similar but Logically Irrelevant Contexts (SSLI), and Counterfactual Contexts (CC). Through systematic evaluation of 13 widely used rerankers on five families of LLMs, we provide detailed insights into their relative strengths and limitations. To the best of our knowledge, SciRerankBench is the first benchmark specifically developed to evaluate rerankers within RAG-LLMs, which provides valuable observations and guidance for their future development.

科学问答RAG重排器评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。