企业级RAG系统在复杂场景下指令遵守率暴跌,暴露出严重可靠性问题。
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

- 构建983个专家验证样本,模拟检索噪声、知识缺失和事实冲突三种真实故障模式
- 13个主流大模型在多约束条件下仅26.8%响应全部达标,存在57点协同差距
- 揭示推理增强也无法克服知识缺口与事实冲突,适合部署决策者和RAG研发者参考
企业级RAG部署面临关键可靠性缺口:尽管大语言模型对单个约束的满足率达80%,但仅有26.8%的回复能同时满足所有要求,暴露出57个百分点的协同差距。现有基准假设干净检索与简单查询,无法反映生产环境中噪声文档与多维约束共存的真实情况。我们提出EnterpriseRAG,包含6个领域共983个专家验证样本,系统模拟三种此前研究中缺失的故障模式:检索噪声、知识缺口和事实冲突,并结合复杂指令进行评估。对13个最先进的大语言模型测试显示,指令遵守率出现严重坍塌,高单约束满足率掩盖了低整体合规性。关键发现表明,在知识缺口与事实冲突下,即使采用增强推理仍存在深层障碍,说明生产级RAG需引入显式的上下文感知协议与校准判断机制。EnterpriseRAG为衡量和弥合这些差距提供了可复现的基础,直接支持企业级RAG系统的部署决策。论文发表后将公开基准与评估框架。
原文摘要 · Abstract (English)
Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。