无需参考答案,全面评估大模型问答系统性能
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering
- 设计六项互相关联的无参考指标,实现端到端评估
- 可定位问答系统中需优化的具体组件环节
- 适合开发者持续监控和改进RAG应用
我们提出THELMA(基于任务的大型语言模型应用综合评估),一个针对基于RAG(检索增强生成)的问答应用的无参考评估框架。THELMA包含六项专门设计的相互关联指标,用于对RAG问答应用进行全面、细粒度的评估。该框架使开发者和应用管理者能够在无需标注数据或参考答案的情况下,评估、监控并改进完整的RAG问答流水线。我们还揭示了所提THELMA指标之间的相互作用关系,这些关系可用于识别需要改进的特定RAG组件。
原文摘要 · Abstract (English)
We propose THELMA (Task Based Holistic Evaluation of Large Language Model Applications), a reference free framework for RAG (Retrieval Augmented generation) based question answering (QA) applications. THELMA consist of six interdependent metrics specifically designed for holistic, fine grained evaluation of RAG QA applications. THELMA framework helps developers and application owners evaluate, monitor and improve end to end RAG QA pipelines without requiring labelled sources or reference responses.We also present our findings on the interplay of the proposed THELMA metrics, which can be interpreted to identify the specific RAG component needing improvement in QA applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。