用贝叶斯模型统一评估RAG系统,揭示隐藏的错误传播机制
The RAT: A Unified Bayesian Model for RAG Evaluation

- 基于信息流分解检索、拒绝与答案正确性的联合概率模型
- 27种配置对比显示,边缘指标相同但行为差异显著
- 支持人工标注与自动化评估融合,提升小样本下的评估效率
评估检索增强生成(RAG)系统不仅需关注端到端正确性,还需分析各组件间交互及错误传播。本文提出一种贝叶斯评估框架,联合建模检索成功、拒绝行为与答案正确性,按流水线信息流进行因子分解。该模型区分任务成功:用户是否获得正确答案(生成器表现),以及生成器在给定检索结果下是否合理应对。我们在三个数据集、三种检索器和三种生成器的27种RAG配置上应用该框架,发现条件分解揭示了边际指标看似相同时系统间显著的行为差异。进一步分析标注分配问题,表明检索成功标注比任务成功标注更有利于估计策略遵循度,并从信息论角度解释该不对称性。最后,将模型扩展为融合大模型作为评判者标注作为校准噪声观测,使从业者可在统一概率框架内结合有限人工判断与低成本自动化评估。
原文摘要 · Abstract (English)
Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow. The model distinguishes task success. Whether the user received a correct answer (from generator success) and whether the generator behaved appropriately given the retrieval outcome. We apply the framework to 27 RAG configurations across three datasets, three retrievers, and three generators, and show that the conditional decomposition reveals substantial behavioral differences between systems that appear equivalent under marginal metrics. We further analyze the annotation allocation problem, demonstrating that retrieval-success annotations are more informative than task-success annotations for estimating policy adherence, and provide an information-theoretic explanation for this asymmetry. Finally, we extend the model to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited human judgments with cheaper automated assessments within a unified probabilistic model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。