arXiv:2502.20377cs.LGcs.AI2025-02ICML被引 14

按需生成独特文档与问答对,避免模型作弊,更真实评估推理能力。

PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation

  • 每次评测动态生成唯一文档库和问题,杜绝数据泄露
  • 不同难度与规模测试下,前沿大模型表现显著下降
  • 适合需要严谨评估推理与检索能力的研究者

高质量基准测试对评估大语言模型的推理与检索能力至关重要。然而,传统数据集易导致数据泄露,造成性能虚高。为此,我们提出PhantomWiki:一种按需生成独特、事实一致文档语料库及多样化问答对的流水线。不同于已有固定数据集或基于现有数据的方法,PhantomWiki每次评估时动态生成新实例。通过调节问题难度和语料规模,可分别解耦推理与检索能力。实验发现,即使在前沿大模型上,PhantomWiki仍极具挑战性。我们贡献了一个可扩展、抗数据泄露的评估框架,用于分离评估推理、检索与工具使用能力。代码已开源:https://github.com/kilian-group/phantom-wiki。

原文摘要 · Abstract (English)

High-quality benchmarks are essential for evaluating reasoning and retrieval capabilities of large language models (LLMs). However, curating datasets for this purpose is not a permanent solution as they are prone to data leakage and inflated performance results. To address these challenges, we propose PhantomWiki: a pipeline to generate unique, factually consistent document corpora with diverse question-answer pairs. Unlike prior work, PhantomWiki is neither a fixed dataset, nor is it based on any existing data. Instead, a new PhantomWiki instance is generated on demand for each evaluation. We vary the question difficulty and corpus size to disentangle reasoning and retrieval capabilities respectively, and find that PhantomWiki datasets are surprisingly challenging for frontier LLMs. Thus, we contribute a scalable and data leakage-resistant framework for disentangled evaluation of reasoning, retrieval, and tool-use abilities. Our code is available at https://github.com/kilian-group/phantom-wiki.

评估基准推理能力数据泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。