RankArena统一评估检索与生成系统,支持人机双视角反馈。
RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback
- 融合人类与大模型反馈,支持多种对比评估模式。
- 可捕捉细粒度相关性偏好和标注时间等辅助数据。
- 适合评测检索、重排及RAG系统,也适配训练奖励模型。
由于缺乏可扩展、以用户为中心且多视角的评估工具,检索增强生成(RAG)与文档重排系统的质量评估仍具挑战。我们提出RankArena,一个统一平台,用于比较和分析检索流水线、重排器及RAG系统的性能,采用结构化的人类与大模型反馈,并支持此类反馈的收集。该平台支持多种评估模式:直接重排可视化、盲态成对比较(人工或大模型投票)、监督式手动文档标注,以及端到端 RAG 答案质量评估。通过成对偏好与全列表标注,捕捉细粒度相关性反馈,并附带移动指标、标注时长与质量评分等辅助元数据。平台还集成大模型作为裁判的评估能力,实现模型生成排名与人工标注真值之间的对比。所有交互均以结构化评估数据集形式存储,可用于训练重排器、奖励模型、判断代理或检索策略选择器。平台已公开,地址为 https://rankarena.ngrok.io/,演示视频见 https://youtu.be/jIYAP4PaSSI。
原文摘要 · Abstract (English)
Evaluating the quality of retrieval-augmented generation (RAG) and document reranking systems remains challenging due to the lack of scalable, user-centric, and multi-perspective evaluation tools. We introduce RankArena, a unified platform for comparing and analysing the performance of retrieval pipelines, rerankers, and RAG systems using structured human and LLM-based feedback as well as for collecting such feedback. RankArena supports multiple evaluation modes: direct reranking visualisation, blind pairwise comparisons with human or LLM voting, supervised manual document annotation, and end-to-end RAG answer quality assessment. It captures fine-grained relevance feedback through both pairwise preferences and full-list annotations, along with auxiliary metadata such as movement metrics, annotation time, and quality ratings. The platform also integrates LLM-as-a-judge evaluation, enabling comparison between model-generated rankings and human ground truth annotations. All interactions are stored as structured evaluation datasets that can be used to train rerankers, reward models, judgment agents, or retrieval strategy selectors. Our platform is publicly available at https://rankarena.ngrok.io/, and the Demo video is provided https://youtu.be/jIYAP4PaSSI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。