arXiv:2609.07879cs.AI2026-09

用新指标测大模型是否知道自己的无知,发现表现差异大。

Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

  • 设计行为性指标EHQ,评估模型对知识边界的诚实态度。
  • 14个模型的综合得分在0.31到0.81之间,差距显著。
  • 结果表明传统评测忽略的“无知识别”能力很重要,适合关注模型可信度的研究者。

大型语言模型(LLMs)常表现自信且表达流畅,但它们是否清楚自己的知识边界?为此,我们引入“认知诚实”概念,提出全新评估指标Epistemic Honesty Quotient(EHQ),从两个维度(认知克制与答案准确性校准)量化模型对未知的回应行为。构建了包含3000道题的EHQ-3000基准,涵盖虚构实体、截止后事件、超小众真实问题和情境依赖问题四类。从21个冻结的API接口中筛选出15个完成测试,其中14个进入正式分析(因一处接口截断严重导致不可靠)。结果显示模型间差异显著,复合EHQ得分范围为0.31至0.81,即使在文档相关任务上表现接近满分。克制性指标重叠度高,而答案校准性则模型间不一,且不随克制性同步变化;受限于样本量,结论仍有不确定性。该研究揭示了传统正确率评估无法捕捉的行为差异,并强调数据集构成、服务端行为及置信度提取方式对解释结果至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.

认知诚实大模型评测可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。