为大模型评判RAG系统制定可复现的评估标准,避免虚假性能夸大。
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test

- 固定候选池、预算和生成参数,明确评估条件
- 集群感知推理使4个基线对比中仅1个显著,避免误判
- 适合关注真实性能的RAG研究者与评测人员
检索增强生成(RAG)系统常通过大语言模型(LLM)裁判比较答案优劣。对于多跳RAG,评分结果可能受检索质量、答案长度、词汇重叠或忽略数据聚类的影响,已不仅是建模问题,更是测量问题。本文提出一种最小评估标准:固定前100个候选、证据预算、答案长度上限、生成器与提示模板;要求预注册假设、集群感知推理、可行时进行精确的集群符号翻转检验,并由第二位裁判复现。在计算机科学/机器学习(CS/ML)与材料科学的400个多跳问题上测试该协议。原方法下四个语义基线均显著,但采用集群感知后仅一个结果在邦弗朗尼校正下显著。相同预算下,BM25优于纯语义的GADMEC,而词法-语义混合策略在CS/ML中恢复表现,缩小了材料科学中的差距。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) systems are often compared by asking a large language model (LLM) judge which answer is better. For multi-hop RAG, this has become a measurement problem as much as a modeling problem: the same score can reflect retrieval quality, answer length, lexical overlap, or a statistical test that ignores clustered data. We ask what happens when these choices are made explicit. We propose a minimum measurement standard for LLM-as-a-judge comparisons in RAG. The standard fixes the top-100 candidate pool, evidence budget, answer cap, generator, and prompt; it also requires pre-registered hypotheses, cluster-aware inference, an exact cluster sign-flip check when feasible, and second-judge replication. Clustered benchmarks can overstate progress; the field should adopt this standard. We stress-test it with Genetic Algorithm Decoder for Multi-hop Evidence Composition (GADMEC), an evolutionary evidence selector, on 400 multi-hop questions in computer science/machine learning (CS/ML) and Materials Science. The protocol changes the empirical story. A binomial test makes all four semantic-baseline comparisons look significant; cluster-aware inference leaves only one Bonferroni-significant result. BM25 beats pure semantic GADMEC under the same budget, while a lexical-semantic hybrid recovers in CS/ML and narrows the Materials Science gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。