让大模型互相打分,用博弈论方法评估其是否符合人类判断。
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
- 大模型自评互评,通过自我对弈和同行评审代替人工标注。
- 实验显示模型评分与人类偏好部分一致,也存在明显差异。
- 适合关注大模型评估方法、对齐研究的学者和工程师。
现有大语言模型评估方法多依赖固定任务与参考答案,难以捕捉模型行为的复杂性与主观性。本文提出一种自动互评框架,让大模型通过自对弈和同行评审相互打分,并用博弈论投票算法聚合评价结果,系统比对模型排名与人类投票的一致性。实验表明,模型生成的排序在部分情况下与人类偏好吻合,但仍有显著偏差,揭示了互评机制的潜力与局限。本工作首次将互评、博弈论聚合与人类基准验证相结合,为大模型能力评估提供新范式。
原文摘要 · Abstract (English)
Ideal or real - that is the question.In this work, we explore whether principles from game theory can be effectively applied to the evaluation of large language models (LLMs). This inquiry is motivated by the growing inadequacy of conventional evaluation practices, which often rely on fixed-format tasks with reference answers and struggle to capture the nuanced, subjective, and open-ended nature of modern LLM behavior. To address these challenges, we propose a novel alternative: automatic mutual evaluation, where LLMs assess each other's output through self-play and peer review. These peer assessments are then systematically compared with human voting behavior to evaluate their alignment with human judgment. Our framework incorporates game-theoretic voting algorithms to aggregate peer reviews, enabling a principled investigation into whether model-generated rankings reflect human preferences. Empirical results reveal both convergences and divergences between theoretical predictions and human evaluations, offering valuable insights into the promises and limitations of mutual evaluation. To the best of our knowledge, this is the first work to jointly integrate mutual evaluation, game-theoretic aggregation, and human-grounded validation for evaluating the capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。