arXiv:2604.02640cs.CL2026-04中稿 · AAAI被引 1

为真实企业场景设计了RAG评估新框架,解决高分低效难题。

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework

  • 构建四维难度分类体系,系统诊断RAG短板
  • 提出企业级RAG基准,覆盖推理复杂度与可解释性要求
  • 适合落地部署前的RAG系统可靠性评估

企业环境中检索增强生成(RAG)系统的性能评估涉及多维度复合因素,远超简单准确率。这些因素包括推理复杂度、检索难度、文档结构多样性以及对操作可解释性的严格要求。现有学术基准无法系统诊断这些相互关联的挑战,导致模型虽获高分却在实际部署中可靠性不足。为此,本研究提出一个多维度诊断框架,通过定义四轴难度分类体系,并将其整合进企业级RAG基准,以识别系统潜在弱点。

原文摘要 · Abstract (English)

Performance evaluation of Retrieval-Augmented Generation (RAG) systems within enterprise environments is governed by multi-dimensional and composite factors extending far beyond simple final accuracy checks. These factors include reasoning complexity, retrieval difficulty, the diverse structure of documents, and stringent requirements for operational explainability. Existing academic benchmarks fail to systematically diagnose these interlocking challenges, resulting in a critical gap where models achieving high performance scores fail to meet the expected reliability in practical deployment. To bridge this discrepancy, this research proposes a multi-dimensional diagnostic framework by defining a four-axis difficulty taxonomy and integrating it into an enterprise RAG benchmark to diagnose potential system weaknesses.

RAG评估基准企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。