测试智能体检索增强生成系统的多层可靠性,发现单一修复无法解决所有问题。
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
- 构建跨层级可靠性基准,涵盖8个企业领域与9种故障场景。
- 模式归一化可提升0.913的结构漂移成功率,但无法修复权限或会话错误。
- 强调需分层评估,避免误判修复措施为万能解法。
智能体检索增强生成系统可能在表面看似有依据,却在证据、工具合约、授权或会话状态等层面失败。我们提出 LayerRAG-Bench,一个受控的跨层可靠性基准,包含8个企业领域、240项任务、9种故障情景、2种合约模式,以及来自OpenAI、Anthropic和Gemini的九个模型产生的38,880条实时任务记录。模式归一化将结构漂移的成功率从0.000提升至0.913,但过时证据、缺失工具输出、权限被拒和错误会话上下文仍无法通过该方法恢复。仅评估有据性也会在过时或错误会话证据下产生大量假阳性结果。这些结果支持分层评估原则:可靠性干预应仅对其目标层级负责,不应被误认为通用解决方案。
原文摘要 · Abstract (English)
Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are not recovered by schema normalization. Groundedness-only evaluation also produces substantial false positives under stale and wrong-session evidence. These results support a layer-specific evaluation principle: a reliability intervention should be credited for repairing its target layer without being mistaken for a universal fix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。