用棋类对战评估大模型综合能力,发现其抗压性强但策略稳定性不足。
Who is a Better Player: LLM against LLM
- 设计棋类对战平台,通过循环赛评估大模型表现
- 多数大模型在对战中保持乐观,适应高压环境优于人类
- 发现模型策略存在周期性波动,需深入分析稳定性问题
对抗性棋类游戏作为战略推理与智能的典范领域,长期既是竞技活动,也是评估人工智能系统性能的基准。基于此,我们提出一种基于棋类对战的大语言模型(LLM)综合性能评估框架,弥补主流问答式基准对数据依赖的局限。引入「齐镇」平台,支持5种流行棋类游戏,包含20个由大模型驱动的玩家。平台采用埃洛评级系统和新型性能循环图(PLG)量化评估大模型技术能力,同时通过正向情绪得分(PSS)追踪游戏过程中的心理状态。评估以循环赛形式进行,实现玩家间的系统性对比。实验表明,尽管技术差异存在,多数大模型对胜负均持乐观态度,表现出比人类更强的高压力对抗环境适应性。然而,性能循环图中周期性胜负关系揭示了大模型策略表现的不稳定性,亟待进一步解释与探索。
原文摘要 · Abstract (English)
Adversarial board games, as a paradigmatic domain of strategic reasoning and intelligence, have long served as both a popular competitive activity and a benchmark for evaluating artificial intelligence (AI) systems. Building on this foundation, we propose an adversarial benchmarking framework to assess the comprehensive performance of Large Language Models (LLMs) through board games competition, compensating the limitation of data dependency of the mainstream Question-and-Answer (Q&A) based benchmark method. We introduce Qi Town, a specialized evaluation platform that supports 5 widely played games and involves 20 LLM-driven players. The platform employs both the Elo rating system and a novel Performance Loop Graph (PLG) to quantitatively evaluate the technical capabilities of LLMs, while also capturing Positive Sentiment Score (PSS) throughout gameplay to assess mental fitness. The evaluation is structured as a round-robin tournament, enabling systematic comparison across players. Experimental results indicate that, despite technical differences, most LLMs remain optimistic about winning and losing, demonstrating greater adaptability to high-stress adversarial environments than humans. On the other hand, the complex relationship between cyclic wins and losses in PLGs exposes the instability of LLMs' skill play during games, warranting further explanation and exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。