arXiv:2511.13029cs.CLcs.AI2025-11被引 11

测试大模型跨领域知识可靠性,发现多数模型仍易编造事实。

AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models

  • 构建6000题基准,评估模型事实记忆与不确定性识别能力
  • 仅有3款模型得分超0,最高为Claude 4.1 Opus(4.8分)
  • 不同模型在各领域表现差异大,需按任务选型

现有语言模型评估主要关注通用能力,但跨领域可靠使用需具备事实准确性与知识缺口认知。我们提出AA-Omniscience基准,涵盖6000个问题,来源为权威学术与产业资料,覆盖6个领域的42个经济相关主题。评估采用奥米尼西恩斯指数(-100至100),同时惩罚幻觉与奖励不确定时的回避,0分代表正确与错误回答数量相等。评估显示,仅3款模型得分高于0,其中Claude 4.1 Opus得分最高(4.8)。结果揭示前沿模型普遍存在事实性与校准性缺陷。模型表现随领域变化明显,三家研究实验室的模型分别在不同领域领先。这表明,在知识关键任务中应根据具体需求选择模型,而非仅依赖总体性能。

原文摘要 · Abstract (English)

Existing language model evaluations primarily measure general capabilities, yet reliable use of these models across a range of domains demands factual accuracy and recognition of knowledge gaps. We introduce AA-Omniscience, a benchmark designed to measure both factual recall and knowledge calibration across 6,000 questions. Questions are derived from authoritative academic and industry sources, and cover 42 economically relevant topics within six different domains. The evaluation measures a model's Omniscience Index, a bounded metric (-100 to 100) measuring factual recall that jointly penalizes hallucinations and rewards abstention when uncertain, with 0 equating to a model that answers questions correctly as much as it does incorrectly. Among evaluated models, Claude 4.1 Opus attains the highest score (4.8), making it one of only three models to score above zero. These results reveal persistent factuality and calibration weaknesses across frontier models. Performance also varies by domain, with the models from three different research labs leading across the six domains. This performance variability suggests models should be chosen according to the demands of the use case rather than general performance for tasks where knowledge is important.

大模型评估知识可靠性事实性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。