RAG在复杂问答中表现不佳,通过结构化输出可显著超越长上下文模型。
Evaluation of retrieval-based QA on QUEST-LOFT
- 引入结构化输出格式,融合推理与证据链
- 在QUEST-LOFT上优化后性能超越长上下文模型
- 适合需要高可信度答案的开放域问答场景
尽管检索增强生成(RAG)在学术界和工业界广受欢迎,但在信息分散于多文档或需复杂推理的问答任务中仍表现不佳。近期的LOFT研究显示,长上下文语言模型同样面临此挑战,尤其在QUEST基准上存在巨大提升空间。本文深入分析QUEST-LOFT表现差的原因,基于全面的人工评估发布更新数据,并证明结合结构化输出(包含推理与证据)的RAG方法,在经答案再验证后可显著优于长上下文模型。
原文摘要 · Abstract (English)
Despite the popularity of retrieval-augmented generation (RAG) as a solution for grounded QA in both academia and industry, current RAG methods struggle with questions where the necessary information is distributed across many documents or where retrieval needs to be combined with complex reasoning. Recently, the LOFT study has shown that this limitation also applies to approaches based on long-context language models, with the QUEST benchmark exhibiting particularly large headroom. In this paper, we provide an in-depth analysis of the factors contributing to the poor performance on QUEST-LOFT, publish updated numbers based on a thorough human evaluation, and demonstrate that RAG can be optimized to significantly outperform long-context approaches when combined with a structured output format containing reasoning and evidence, optionally followed by answer re-verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。