arXiv:2510.10930cs.CLcs.AI2025-10被引 5

测试大模型对游戏的评价能力,发现越聪明的模型反而越不像人。

Evaluating Language Models' Evaluations of Games

  • 用新框架评估模型对游戏的评价,对比人类与计算代理。
  • 推理模型比普通语言模型更像人,但接近最优时反而偏离人类判断。
  • 评价游戏趣味性更不稳定,体现其主观难量化,适合研究认知偏差者看。

推理不仅是解题,更是判断什么问题值得解决。传统AI评估多关注模型下棋等任务表现。本文提出新范式:评估模型自身对游戏的评价能力。我们构建了包含100多个原创棋类游戏和450多条人工评分的大规模数据集,比较现代语言与推理模型、人类及符号计算代理在评估游戏收益(公平性)和趣味性上的表现。这两类问题分别对应评估复杂度与量化难度两个维度。结果表明,推理模型整体上比非推理模型更贴近人类评价;但存在非单调关系——当模型越接近博弈论最优时,与人类数据的契合度反而下降。此外,在趣味性评估上模型表现更不平滑,印证其主观性与量化难度。各类任务中,推理模型资源使用波动极大,凸显需引入更合理的元推理机制。

原文摘要 · Abstract (English)

Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily focused on problem solving, historically by studying how models play games such as chess and Go. In this paper, we advocate for a new paradigm that assesses AI systems' evaluation of games. First, we introduce a formalism for evaluating such evaluations. We then leverage a large-scale dataset of over 100 novel board games and over 450 human judgments to compare evaluations produced by modern language and reasoning models against those of people and symbolic computational agents. We consider two kinds of evaluative queries: assessing the payoff (or fairness) and the funness of games. These queries span two dimensions relevant to the design of evaluations of AI evaluations: how complex a query is to compute and how difficult a query is to quantify. Our results show that reasoning models are generally more aligned to people in their evaluations of games than non-reasoning language models. However, we observe a non-monotonic relationship: as models get closer to game-theoretic optimal, their fit to human data weakens. We also observe more "jaggedness" across models for assessing funness, in line with the greater difficulty of quantifying this query. Across queries and games, reasoning models show highly variable and unpredictable resource usage when assessing queries, pointing to the importance of imbuing more resource-rational meta-reasoning in language and reasoning models.

游戏评估大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。