首个系统评估语言模型用外源知识能力的基准,揭示现有方法在真实场景下的短板。
CUB: Benchmarking Context Utilisation Techniques for Language Models
- 构建CUB基准,模拟多种噪声上下文测试语言模型
- 7种主流方法在真实数据上表现普遍下降,平均性能差15%以上
- 提醒研究者警惕合成数据上的虚假高分,适合做RAG评估的研究者参考
外部知识对问答、事实核查等知识密集型任务至关重要。然而,语言模型可能忽略与过时参数记忆冲突的信息,或被无关上下文干扰。尽管近期提出多种上下文利用调控技术(CMTs)以缓解此问题,但缺乏系统性比较。本文提出CUB(Context Utilisation Benchmark)——首个专为检索增强生成(RAG)中多样化噪声上下文设计的综合性基准。基于该基准,我们对七种代表主流类别的先进CMTs进行了最全面评估,覆盖三个不同数据集和任务,应用于11个语言模型。结果揭示当前评估实践存在明显缺陷,凸显全面测试的必要性。发现多数现有CMT难以应对真实RAG场景中的全谱上下文类型。此外,许多CMT在简单合成数据集上表现优异,但在包含自然样本的真实数据集上性能显著下降。
原文摘要 · Abstract (English)
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。