arXiv:2605.26937cs.CLcs.AI2026-05

用开放提问评估大模型真实知识表达,打破固定问题的偏见。

Beyond Questions: Evaluating LLM's Knowledge Expression

论文配图:Beyond Questions: Evaluating LLM's Knowledge Expression
图 1 · 摘自论文原文
  • 不设固定问题,让模型自由输出已知信息
  • 覆盖1万实体,验证模型表达内容的准确性
  • 适合研究模型真实知识能力与表达机制的人

大型语言模型中的参数化知识是其成功的关键,但至今仍缺乏深入理解。现有知识评测通常依赖预设问题(如“马丁·路德·金的出生日期是什么?”),仅评估评测设计者选定的知识点,存在显著的可用性偏差。本文提出开放知识评估新范式,不再使用窄义问题,而是通过开放式引导提示(如“告诉我你对马丁·路德·金了解的一切”)评估模型主动呈现的知识。这一转变将重点从预设答案检索转向刻画模型自然表达的知识特征。我们构建了包含10,000个实体及参考语料库的BeQu(Beyond Questions)基准,用于陈述句验证。利用BeQu,我们评估了多种语言模型,并分析了推理努力、模型规模、提示格式和知识领域的影响。数据与排行榜已在项目GitHub及基准网站公开。

原文摘要 · Abstract (English)

Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge benchmarks typically rely on predefined questions (e.g., "What is the birth date of M.L. King?"), evaluating only knowledge that benchmark designers explicitly choose to query, a problematic availability bias. In this paper, we introduce open knowledge evaluation, a new paradigm for LLM knowledge expression benchmarking. Instead of asking narrow questions, it evaluates models on the knowledge they choose to surface in response to open-ended elicitation prompts (e.g., "Tell me everything you know about M.L. King"). This shifts the focus from predefined answer retrieval toward characterizing the knowledge models naturally express. We instantiate this paradigm with BeQu (Beyond Questions), a benchmark of 10,000 entities paired with reference corpora for statement verification. Using BeQu, we evaluate a broad range of language models and analyze the effects of reasoning effort, model scale, prompt format, and knowledge domain. Data and leaderboard are available on this work's GitHub repository and at the benchmark's website.

大模型评测知识表达开放问答语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。