arXiv:2411.14486cs.CLcs.AI2024-11

测试大模型在675个无解问题中承认无知的能力,揭示其认知边界。

The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz

  • 构建包含675个不可解难题的数据集,评估模型是否能承认未知。
  • 最优模型在哲学等领域的认错率仅达62%-68%,越难问题越易承认无知。
  • 适用于评估AGI能力,尤其关注模型对自身知识边界的认知。

本研究提出一种新型评估框架,用于测试大语言模型(LLMs)在675个根本不可解的问题上承认不确定性的能力。通过精心筛选的研究生级重大挑战问题(答案故意无法获得),评估了12个最先进的LLM(包括开源与闭源模型)在不生成看似合理但错误回答时承认无知的倾向。最佳模型在生物学、哲学和数学等领域承认问题无解的准确率介于62%至68%之间。观察到问题难度与模型准确率呈负相关:GPT-4在更困难的问题上承认不确定性比例更高(35.8%),而在简单问题上仅为20.0%。这表明模型在看似可解的问题上更倾向于猜测。不同题型表现差异显著,模型在发明类和NP-hard问题中难以承认无知,而在哲学与心理学问题上表现较好。该研究为人工智能通用性(AGI)评估提供了实证依据,强调识别不确定性是未来机器智能评价的关键维度,推动了对当前大模型认知局限性的理解,并为训练架构与评估方法的改进指明方向。

原文摘要 · Abstract (English)

This research introduces a novel evaluation framework designed to assess large language models' (LLMs) ability to acknowledge uncertainty on 675 fundamentally unsolvable problems. Using a curated dataset of graduate-level grand challenge questions with intentionally unknowable answers, we evaluated twelve state-of-the-art LLMs, including both open and closed-source models, on their propensity to admit ignorance rather than generate plausible but incorrect responses. The best models scored in 62-68% accuracy ranges for admitting the problem solution was unknown in fields ranging from biology to philosophy and mathematics. We observed an inverse relationship between problem difficulty and model accuracy, with GPT-4 demonstrating higher rates of uncertainty acknowledgment on more challenging problems (35.8%) compared to simpler ones (20.0%). This pattern indicates that models may be more prone to generate speculative answers when problems appear more tractable. The study also revealed significant variations across problem categories, with models showing difficulty in acknowledging uncertainty in invention and NP-hard problems while performing relatively better on philosophical and psychological challenges. These results contribute to the growing body of research on artificial general intelligence (AGI) assessment by highlighting the importance of uncertainty recognition as a critical component of future machine intelligence evaluation. This impossibility test thus extends previous theoretical frameworks for universal intelligence testing by providing empirical evidence of current limitations in LLMs' ability to recognize their own knowledge boundaries, suggesting new directions for improving model training architectures and evaluation approaches.

大模型评测不确定性识别AGI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。