arXiv:2504.07803cs.CLcs.AI2025-04被引 4

为RAG系统评估提供一站式黑盒测试框架,助力真实场景对比。

A System for Comprehensive Assessment of RAG Frameworks

  • 构建模块化评估框架,支持多种部署配置的自动化测试
  • 通过API接口实现跨向量库与LLM策略的性能对比报告
  • 兼顾回答连贯性等实用因素,适合研究与工业界评估

检索增强生成(RAG)已成为提升大语言模型事实准确性和上下文相关性的标准范式。然而,现有评估框架难以在真实部署场景中提供全面的黑盒评估。为此,我们提出SCARF(RAG框架综合评估系统),一个模块化、灵活的评估框架,可系统性地基准测试已部署的RAG应用。SCARF提供端到端的黑盒评估方法,支持多种部署配置,自动测试不同向量数据库和大模型服务策略,生成详细性能报告。同时,框架集成响应连贯性等实际考量,具备可扩展性与适应性,适用于研究人员和行业从业者对RAG应用的评估。通过REST API接口,我们展示了其在真实场景中的应用,验证了其在评估不同RAG框架及配置时的灵活性。SCARF开源于GitHub。

原文摘要 · Abstract (English)

Retrieval Augmented Generation (RAG) has emerged as a standard paradigm for enhancing the factual accuracy and contextual relevance of Large Language Models (LLMs) by integrating retrieval mechanisms. However, existing evaluation frameworks fail to provide a holistic black-box approach to assessing RAG systems, especially in real-world deployment scenarios. To address this gap, we introduce SCARF (System for Comprehensive Assessment of RAG Frameworks), a modular and flexible evaluation framework designed to benchmark deployed RAG applications systematically. SCARF provides an end-to-end, black-box evaluation methodology, enabling a limited-effort comparison across diverse RAG frameworks. Our framework supports multiple deployment configurations and facilitates automated testing across vector databases and LLM serving strategies, producing a detailed performance report. Moreover, SCARF integrates practical considerations such as response coherence, providing a scalable and adaptable solution for researchers and industry professionals evaluating RAG applications. Using the REST APIs interface, we demonstrate how SCARF can be applied to real-world scenarios, showcasing its flexibility in assessing different RAG frameworks and configurations. SCARF is available at GitHub repository.

RAG评估大模型自动化测试黑盒评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。