arXiv:2504.12972cs.CL2025-04被引 1

提出自动估算混合检索增强摘要最优上下文长度的方法。

Estimating Optimal Context Length for Hybrid Retrieval-augmented Multi-document Summarization

  • 基于检索器、摘要器和数据集特性,估算最优检索长度。
  • 通过大模型生成银标准参考,验证不同配置下的最佳上下文长度。
  • 适用于各类大模型,尤其对超长上下文模型效果显著。

近期语言模型在长上下文推理方面取得进展,推动了大规模多文档摘要的应用。然而,已有研究表明,这些长上下文模型并未有效利用其宣称的上下文窗口。为此,检索增强系统成为高效且有效的替代方案,但其性能高度依赖于检索上下文长度的选择。本文提出一种混合方法,结合检索增强系统与最新语言模型支持的长上下文窗口。首先,根据检索器、摘要器和数据集特性,估计最优检索长度;在数据集的随机子集上,使用一组大模型生成银标准参考文本;利用这些银标准参考,评估特定RAG系统配置下的最优上下文长度。在多文档摘要任务上的实验表明,该方法在不同模型类别和规模下均有效。与RULER和HELMET等强基线长上下文基准对比,本方法表现更优,并展现出对超长上下文语言模型的有效性及向新模型类别的良好泛化能力。

原文摘要 · Abstract (English)

Recent advances in long-context reasoning abilities of language models led to interesting applications in large-scale multi-document summarization. However, prior work has shown that these long-context models are not effective at their claimed context windows. To this end, retrieval-augmented systems provide an efficient and effective alternative. However, their performance can be highly sensitive to the choice of retrieval context length. In this work, we present a hybrid method that combines retrieval-augmented systems with long-context windows supported by recent language models. Our method first estimates the optimal retrieval length as a function of the retriever, summarizer, and dataset. On a randomly sampled subset of the dataset, we use a panel of LLMs to generate a pool of silver references. We use these silver references to estimate the optimal context length for a given RAG system configuration. Our results on the multi-document summarization task showcase the effectiveness of our method across model classes and sizes. We compare against length estimates from strong long-context benchmarks such as RULER and HELMET. Our analysis also highlights the effectiveness of our estimation method for very long-context LMs and its generalization to new classes of LMs.

摘要生成检索增强上下文长度大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。