arXiv:2602.01590cs.CL2026-02被引 1

用权威维基百科文章评测智能研究代理,发现现有系统差距显著。

Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles

  • 以最新优质维基文章为专家级参考,构建实时评测基准。
  • 39项细粒度写作标准+事实可验证性指标,全面评估生成质量。
  • 揭示当前智能研究代理与人类专家水平的明显差距,适合研究者使用。

深度研究代理(DRAs)在自主信息检索与报告生成方面展现出强大能力,有望协助人类完成复杂研究任务。现有评估框架多依赖大模型生成的参考或评估维度,虽具可扩展性,但缺乏专家验证内容的可靠性,难以提供客观、精细的评估。为此,我们提出Wiki Live Challenge(WLC),利用最新维基百科优质文章(GAs)作为专家级参考。维基百科对中立性、全面性和可验证性的严格标准,使GAs成为对DRAs的极高挑战。我们整理了100篇近期优质文章,提出Wiki Eval评估框架,包含39项细粒度写作质量标准及严格的事实可验证性指标。对多种DRAs系统的大量实验显示,当前代理与人类专家级维基文章间存在显著差距,验证了WLC在推动代理研究方面的有效性。基准已开源:https://github.com/WangShao2000/Wiki_Live_Challenge。

原文摘要 · Abstract (English)

Deep Research Agents (DRAs) have demonstrated remarkable capabilities in autonomous information retrieval and report generation, showing great potential to assist humans in complex research tasks. Current evaluation frameworks primarily rely on LLM-generated references or LLM-derived evaluation dimensions. While these approaches offer scalability, they often lack the reliability of expert-verified content and struggle to provide objective, fine-grained assessments of critical dimensions. To bridge this gap, we introduce Wiki Live Challenge (WLC), a live benchmark that leverages the newest Wikipedia Good Articles (GAs) as expert-level references. Wikipedia's strict standards for neutrality, comprehensiveness, and verifiability serve as a great challenge for DRAs, with GAs representing the pinnacle of which. We curate a dataset of 100 recent Good Articles and propose Wiki Eval, a comprehensive evaluation framework comprising a fine-grained evaluation method with 39 criteria for writing quality and rigorous metrics for factual verifiability. Extensive experiments on various DRA systems demonstrate a significant gap between current DRAs and human expert-level Wikipedia articles, validating the effectiveness of WLC in advancing agent research. We release our benchmark at https://github.com/WangShao2000/Wiki_Live_Challenge

研究代理评估基准维基百科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。