用反向生成文档方法,实现无标注、抗污染的AI知识检索评估。
Scalable and Reliable Evaluation of AI Knowledge Retrieval Systems: RIKER and the Coherent Simulated Universe
- 从已知真实答案反向生成文档,实现确定性评分。
- 33个模型测试显示,超过32K上下文后性能显著下降。
- 揭示事实查找与幻觉抵抗是两种独立能力,适合研究者参考。
评估知识系统(如大模型、RAG、知识图谱等)面临根本挑战:静态基准易被污染,基于大模型的评价者存在系统性偏差,而真实答案提取需昂贵的人工标注。我们提出RIKER(检索智能与知识提取评分),一种基于范式反转的基准与可复现方法——从已知真实答案生成文档,而非从文档中提取真实答案。该方法实现确定性评分,无需人工标注或参考模型,且通过可再生语料库实现抗污染能力。使用超过210亿个标记对33个模型的评估表明:上下文长度宣称常超出实际可用容量,超过32K标记后性能显著下降;跨文档聚合远比单文档提取困难;事实发现能力与幻觉抵抗是两个独立维度——能准确找出现有事实的模型仍可能编造不存在的事实。除具体基准外,我们还提供一种通用方法论,适用于任何可从结构化真实答案生成合成文档的领域。
原文摘要 · Abstract (English)
Evaluating knowledge systems (LLMs, RAG, knowledge graphs, etc) faces fundamental challenges: static benchmarks are vulnerable to contamination, LLM-based judges exhibit systematic biases, and ground truth extraction requires expensive human annotation. We present RIKER (Retrieval Intelligence and Knowledge Extraction Rating), both a benchmark and a replicable methodology based on paradigm inversion - generating documents from known ground truth rather than extracting ground truth from documents. This approach enables deterministic scoring and scalable evaluation without human annotation or reference models, and contamination resistance through regenerable corpora. Our evaluation of 33 models using over 21 billion tokens reveals that context length claims frequently exceed usable capacity, with significant degradation beyond 32K tokens; cross-document aggregation proves substantially harder than single-document extraction; and grounding ability and hallucination resistance are distinct capabilities - models excelling at finding facts that exist may still fabricate facts that do not. Beyond the specific benchmark, we contribute a domain-agnostic methodology for constructing scalable and contamination-resistant evaluations wherever synthetic documents can be generated from structured ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。