用真人多样性评估大模型常识,发现小模型反而更接近人类平均水准
A large-scale evaluation of commonsense knowledge in humans and large language models
- 把大模型当作人群样本,比对其与真实人类的常识判断一致性
- 多数大模型个体表现低于人类中位数,大模型间共识度不高
- 开源小模型表现优于闭源大模型,适合研究人类认知差异
常识是人工智能的重要组成部分,现有评估主要依赖人工设定的标准标签。这一做法隐含假设:人类常识具有一致性。然而近期研究表明,人们对常识的认知存在显著差异——某人认为理所当然的事,对另一人可能并不显然。本文提出一种新方法,通过测量大语言模型(LLMs)的判断与真实人类群体的一致性,来评估其常识能力。结果显示:当将大模型视为独立受访者时,大多数模型的常识能力低于人类中位数;当模拟虚拟人群时,大模型与真实人类在判断一致性的相关性仅达中等水平。有趣的是,小型开放权重模型在多项指标上反而优于大型封闭模型。该评估框架将常识与文化背景关联,支持了将AI适配于不同社会知识体系的必要性。
原文摘要 · Abstract (English)
Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implicit, assumption of these labels is that they accurately capture what any human would think, effectively treating human common sense as homogeneous. However, recent empirical work has shown that humans vary enormously in what they consider commonsensical; thus what appears self-evident to one benchmark designer may not be so to another. Here, we propose a method for assessing commonsense knowledge in AI, specifically in large language models (LLMs), that incorporates empirically observed heterogeneity among humans by measuring the correspondence between a model's judgment and that of a human population. We first find that, when treated as independent survey respondents, most LLMs remain below the human median in their individual commonsense competence. Second, when used as simulators of a hypothetical population, LLMs correlate with real humans only modestly in the extent to which they agree on the same set of statements. In both cases, smaller, open-weight models are surprisingly more competitive than larger, proprietary frontier models. Our evaluation framework, which ties commonsense knowledge to its cultural basis, contributes to the growing call for adapting AI models to human collectivities that possess different, often incompatible, social stocks of knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。