arXiv:2508.10428cs.LG2025-08被引 2

构建复杂决策评估框架,让大模型在星际争霸中自我进化

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

  • 设计支持全种族、低级动作与文本观测的完整评测环境
  • 通过自纠错机制提升策略规划能力,显著优于传统方法
  • 适合研究通用智能体与自进化AI的开发者参考

评估大语言模型在复杂决策任务中的表现,对推动AI的战略规划与实时适应能力至关重要。现有星际争霸II的评测基准未能充分涵盖游戏的全部复杂性,如完整游戏上下文、多样化动作空间及所有可玩种族。为此,我们提出SC2Arena,一个支持所有可玩种族、低级动作空间,并优化文本观测以应对空间推理挑战的完整评测基准。同时,我们引入StarEvolve,一种分层式自改进框架,整合战略规划与战术执行,通过高质游戏数据持续微调实现迭代自修正。其核心包含规划-执行-验证结构,以及用于筛选高质量训练样本的评分系统。基于SC2Arena的全面分析揭示了发展通用智能体的新洞见,实验表明StarEvolve在战略规划上表现更优。代码、环境与算法均已公开。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) in complex decision-making is essential for advancing AI's ability for strategic planning and real-time adaptation. However, existing benchmarks for tasks like StarCraft II fail to capture the game's full complexity, such as its complete game context, diverse action spaces, and all playable races. To address this gap, we present SC2Arena, a benchmark that fully supports all playable races, low-level action spaces, and optimizes text-based observations to tackle spatial reasoning challenges. Complementing this, we introduce StarEvolve, a hierarchical framework that integrates strategic planning with tactical execution, featuring iterative self-correction and continuous improvement via fine-tuning on high-quality gameplay data. Its key components include a Planner-Executor-Verifier structure to break down gameplay, and a scoring system for selecting high-quality training samples. Comprehensive analysis using SC2Arena provides valuable insights into developing generalist agents that were not possible with previous benchmarks. Experimental results also demonstrate that our proposed StarEvolve achieves superior performance in strategic planning. Our code, environment, and algorithms are publicly available.

大模型评估自进化星际争霸决策规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。