构建动态知识评测框架,让大模型评估跟得上实时信息变化。
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
- 用自动化代理生成每日更新的新闻类评测数据
- 小模型加检索后性能接近大模型,差距缩小40%以上
- 适合关注模型真实认知能力与检索增强效果的研究者
知识密集型问答是大语言模型的核心能力,传统评测依赖维基百科等静态数据源,难以反映动态世界的知识演进。为此,我们提出OKBench,一个全自动、按需生成的开放知识评测框架。聚焦每日更新的新闻领域,该框架通过智能代理实现评测数据的自动采集、生成、验证与分发。它降低了评测门槛,促进对检索增强方法的全面评估,并减少与预训练数据的重叠。我们在多种开源及商用大模型上进行测试,涵盖不同规模和配置,对比有无检索的情况。结果表明,面对新知识时模型行为差异显著,且检索能有效缩小小模型与大模型间的性能差距(提升超40%)。研究强调了在动态知识基准上评估大模型的重要性。
原文摘要 · Abstract (English)
Knowledge-intensive question answering is central to large language models (LLMs) and is typically assessed using static benchmarks derived from sources like Wikipedia and textbooks. However, these benchmarks fail to capture evolving knowledge in a dynamic world, and centralized curation struggles to keep pace with rapid LLM advancements. To address these drawbacks, we propose Open Knowledge Bench (OKBench), a fully automated framework for generating high-quality, dynamic knowledge benchmarks on demand. Focusing on the news domain where knowledge updates daily, OKBench is an agentic framework that automates the sourcing, creation, validation, and distribution of benchmarks. Our approach democratizes benchmark creation and facilitates thorough evaluation of retrieval-augmented methods by reducing overlap with pretraining data. We evaluate our framework on a wide range open-source and proprietary LLMs of various sizes and configurations, both with and without retrieval over freshly generated knowledge. Our results reveal distinct model behaviors when confronted with new information and highlight how retrieval narrows the performance gap between small and large models. These findings underscore the importance of evaluating LLMs on evolving knowledge benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。