让大模型互相出题考对方,用动态对抗评估模型真实能力。
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

- 模型轮流出题考对手,根据对方答错与否获得奖励。
- 在10个前沿大模型上构建稳定排名,识别出各自知识盲区。
- 适合开发者检测模型缺陷,不依赖人类偏好且抗数据污染。
评估前沿大模型面临挑战:静态基准存在数据污染和饱和问题,导致用户难以区分顶尖模型,开发者也难以发现特定失效模式;而人类偏好主观性强。本文提出核心问题:大模型是否知道其他模型不知道的内容?能否利用这种特性进行评估?我们提出「LivingArena」——一种自动化、抗污染的评估框架。在该框架中,模型轮流扮演出题者与答题者,出题者旨在设计对手无法正确回答的问题,若对手答错则获奖励;答题者则在正确回答时获得奖励。为确保问题有客观可验证答案,由强模型组成的裁判组进行验证,若问题无法验证则惩罚出题者。对十款前沿大模型的评估结果显示,系统生成了稳定的Elo排名。行为分析表明,模型能主动识别并攻击对手的认知边界,自对战和锦标赛日志显示它们会聚焦并强化对手的薄弱维度。该方法不仅衡量知识记忆,更评估事实严谨性及探测对手弱点的高阶能力,与人类偏好相关性弱,提供了一种可扩展、低成本的持续评估方案。
原文摘要 · Abstract (English)
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。