测试大模型在对抗环境中的策略与快速决策能力,发现推理强的模型不一定反应快。
Beyond Scaling: Assessing Strategic Reasoning and Rapid Decision-Making Capability of LLMs in Zero-sum Environments
- 设计多智能体对抗框架,模拟回合制与实时赛况下的策略博弈
- 发现推理强的模型在实时对战中表现反而不如响应快的模型
- 适合研究模型策略与执行速度权衡的AI研究人员参考
大型语言模型在静态推理任务上表现优异,但在对抗性、时间敏感的交互环境中作为智能体的有效性仍不清晰。现有评估多将推理视为一次性能力,忽视了对手感知、时间约束和高压执行等挑战。本文提出战略战术推理(STAR)基准,通过1v1零和竞争互动,将推理视为迭代适应的决策过程。该框架支持回合制与实时两种模式,可统一分析长期战略规划与快速战术执行。基于模块化架构与标准化API,STAR支持可复现评估与灵活任务定制。为超越胜负二元结果,引入战略评估套件,衡量战略行为质量如执行效率与结果稳定性。成对评估显示显著的策略-执行差距:虽推理密集型模型在回合制中占优,但其推理延迟使其在实时场景中表现更差,而更快的指令微调模型反而胜出。结果表明,交互环境中的战略智能不仅依赖推理深度,还取决于将计划转化为及时行动的能力,使STAR成为研究这一权衡的可靠基准。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved strong performance on static reasoning benchmarks, yet their effectiveness as interactive agents operating in adversarial, time-sensitive environments remains poorly understood. Existing evaluations largely treat reasoning as a single-shot capability, overlooking the challenges of opponent-aware decision-making, temporal constraints, and execution under pressure. This paper introduces Strategic Tactical Agent Reasoning (STAR) Benchmark, a multi-agent evaluation framework that assesses LLMs through 1v1 zero-sum competitive interactions, framing reasoning as an iterative, adaptive decision-making process. STAR supports both turn-based and real-time settings, enabling controlled analysis of long-horizon strategic planning and fast-paced tactical execution within a unified environment. Built on a modular architecture with a standardized API and fully implemented execution engine, STAR facilitates reproducible evaluation and flexible task customization. To move beyond binary win-loss outcomes, we introduce a Strategic Evaluation Suite that assesses not only competitive success but also the quality of strategic behavior, such as execution efficiency and outcome stability. Extensive pairwise evaluations reveal a pronounced strategy-execution gap: while reasoning-intensive models dominate turn-based settings, their inference latency often leads to inferior performance in real-time scenarios, where faster instruction-tuned models prevail. These results show that strategic intelligence in interactive environments depends not only on reasoning depth, but also on the ability to translate plans into timely actions, positioning STAR as a principled benchmark for studying this trade-off in competitive, dynamic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。