arXiv:2503.06047cs.AIcs.CL2025-03被引 20

构建复杂博弈评估平台,测试大模型在长期策略中的表现。

DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments

  • 设计六类复杂战略游戏,支持多难度定制。
  • 从五个维度评分,精准衡量决策能力。
  • 可追踪策略轨迹,揭示模型深层缺陷。

基于大语言模型(LLM)的智能体正被广泛应用于需要长周期推理、多智能体交互和不确定性决策的复杂战略环境。然而,现有基准大多仅评估单一技能,缺乏环境多样性或依赖宽泛的整体指标。为此,我们提出DSGBench,一个更严谨的战略决策任务评估平台。首先,它包含六种复杂的策略游戏,因其长期性、多维决策需求及任务难度与目标的可调节性,成为理想的测试场景。其次,DSGBench采用细粒度评分体系,从五个具体维度评估决策表现,实现更全面、更科学的评估。此外,平台还集成自动化决策追踪机制,支持对智能体行为模式及策略转折点的深度分析。我们评估了六种主流LLM智能体(含开源与闭源模型),发现其在不同任务中表现出显著差异。通过决策轨迹分析,进一步识别出各类大模型的系统性局限。这些发现为模型选型及未来LLM智能体发展提供了重要参考。

原文摘要 · Abstract (English)

Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making under uncertainty. However, common existing benchmarks either assess isolated skills, lack environmental diversity, or rely on broad overall metrics. To address these issues, we introduce DSGBench, a more rigorous evaluation platform for strategic decision-making tasks. Firstly, it incorporates six complex strategic games which serve as ideal testbeds due to their long-term and multi-dimensional decision-making demands and flexibility in customizing tasks with various difficulty levels and targets. Secondly, DSGBench employs a fine-grained evaluation scoring system which examines the decision-making capabilities by looking into the performance in five specific dimensions, offering a comprehensive assessment in a better-designed fashion. Furthermore, DSGBench also incorporates an automated decision-tracking mechanism which enables in-depth analysis of agent behaviour patterns and the turning points in their strategies. We evaluate six popular LLM agents, including open-source and closed-source models, and observe distinct strengths and limitations among various tasks. Through decision trajectory analysis, we further identify systemic limitations in different LLMs. These findings offer valuable insights for model selection and future LLM-based agent development.

策略博弈大模型评估智能体决策分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。