arXiv:2504.12332cs.CLcs.CY2025-04被引 1

用人类能力框架评估大模型,发现小模型表现可类比人类。

Can the capability of Large Language Models be described by human ability? A Meta Study

  • 用6类11项人类能力指标,评测超80个模型。
  • 参数少于10亿的模型能力可用人类指标描述。
  • 大模型能力随参数量变化显著,且各项能力独立。

大型语言模型(LLMs)的使用者常将其视为具备类人能力的智能体,但其能力与人类的接近程度仍存争议。本文收集了超过80个模型在37个评估基准上的性能数据,这些基准涵盖人类的6种主要能力与11种子能力。通过聚类分析将模型表现分组,并与人类能力分类对比,得出结论:1. 参数少于100亿的某些模型能力确实可用人类能力指标描述;2. 人类中关联性强的能力在模型中几乎不相关;3. 模型能力随参数规模变化显著。

原文摘要 · Abstract (English)

Users of Large Language Models (LLMs) often perceive these models as intelligent entities with human-like capabilities. However, the extent to which LLMs' capabilities truly approximate human abilities remains a topic of debate. In this paper, to characterize the capabilities of LLMs in relation to human capabilities, we collected performance data from over 80 models across 37 evaluation benchmarks. The evaluation benchmarks are categorized into 6 primary abilities and 11 sub-abilities in human aspect. Then, we then clustered the performance rankings into several categories and compared these clustering results with classifications based on human ability aspects. Our findings lead to the following conclusions: 1. We have confirmed that certain capabilities of LLMs with fewer than 10 billion parameters can indeed be described using human ability metrics; 2. While some abilities are considered interrelated in humans, they appear nearly uncorrelated in LLMs; 3. The capabilities possessed by LLMs vary significantly with the parameter scale of the model.

大模型评估人类能力参数规模能力解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。