arXiv:2507.03711cs.CL2025-07中稿 · MAPR 2025

用越南棋类游戏测试大模型多步规划与决策能力

Can LLMs Play Ô Ăn Quan Game? A Study of Multi-Step Planning and Decision Making

  • 设计攻防不同风格的智能体角色,模拟多样化策略
  • 在Ô Ăn Quan游戏中验证Llama系列模型的策略执行效果
  • 揭示大模型在动态博弈中的推理短板与潜在优势

本文通过传统越南棋类游戏Ô Ăn Quan,考察大语言模型(LLMs)的多步规划与决策能力。该游戏包含一系列策略性棋子移动与吃子机制,为评估模型的战略思维提供了独特环境。我们构建了从激进到防守的不同智能体人格,并以Ô Ăn Quan作为测试平台,评估Llama-3.2-3B-Instruct、Llama-3.1-8B-Instruct及Llama-3.3-70B-Instruct等模型在不同策略下的表现。实验旨在理解这些模型如何进行战略决策、规划动作并应对动态游戏状态,结果将揭示其在推理与策略上的优劣势,深化对大模型通用能力的认知。

原文摘要 · Abstract (English)

In this paper, we explore the ability of large language models (LLMs) to plan and make decisions through the lens of the traditional Vietnamese board game, Ô Ăn Quan. This game, which involves a series of strategic token movements and captures, offers a unique environment for evaluating the decision-making and strategic capabilities of LLMs. Specifically, we develop various agent personas, ranging from aggressive to defensive, and employ the Ô Ăn Quan game as a testbed for assessing LLM performance across different strategies. Through experimentation with models like Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, and Llama-3.3-70B-Instruct, we aim to understand how these models execute strategic decision-making, plan moves, and manage dynamic game states. The results will offer insights into the strengths and weaknesses of LLMs in terms of reasoning and strategy, contributing to a deeper understanding of their general capabilities.

大模型推理策略游戏决策规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。