用动态对抗游戏测试大模型真实表现,发现性能随代际跃迁且存在投机取巧现象。
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- 设计六级淘汰制对抗环境,模拟资源受限与信息不对称场景
- 超50个大模型参评,发现同系列模型性能呈代际跃迁,部分模型走捷径赢局
- 揭示静态评测可能被高阶策略污染,动态评估可补足现有框架
当前大语言模型(LLMs)评测面临数据污染的根本挑战,且多假设理想、资源充足环境,忽视模型在压力下的表现。本文提出 extsc{Squid Game},一个资源受限、信息不对称的动态对抗评测环境,通过交互式游戏评估模型在指令遵循、代码、推理、规划和安全对齐等多方面能力。该环境包含六级淘汰机制,评估超过50个大模型,完成迄今规模最大、最全面的通用大模型动态对抗行为研究。结果发现同一模型家族中性能呈现明显的代际跃迁,并观察到部分模型采用推测性捷径获胜,提示静态基准可能存在更高层次的评估范式污染。对比主流基准与 extsc{Squid Game},表明动态评估可作为静态评估的有效补充。
原文摘要 · Abstract (English)
The potential data contamination issue in contemporary large language models (LLMs) benchmarks presents a fundamental challenge to establishing trustworthy evaluation frameworks. Meanwhile, they predominantly assume benign, resource-rich settings, leaving the behavior of LLMs under pressure unexplored. In this paper, we introduce \textsc{Squid Game}, a dynamic and adversarial evaluation environment with resource-constrained and asymmetric information settings elaborated to evaluate LLMs through interactive gameplay against other LLM opponents. Squid Game consists of six elimination-style levels, focusing on multi-faceted abilities, including instruction-following, code, reasoning, planning, and safety alignment. We evaluate over 50 LLMs on Squid Game, presenting the largest behavioral evaluation study of general LLMs on dynamic adversarial scenarios. We observe a clear generational phase transition in performance in the same model lineage and find evidence that some models resort to speculative shortcuts to win the game, indicating the possibility of higher-level evaluation paradigm contamination in static benchmarks. We also compare prominent LLM benchmarks and \textsc{Squid Game}, highlighting that dynamic evaluation can serve as a complementary part for static evaluations. Project page: https://github.com/zijianchen98/LLM_Squid_Game.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。