arXiv:2503.12349cs.AI2025-03被引 22

评测大模型在社交博弈中的策略与推理能力,发现其应对复杂协作仍存短板。

SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?

  • 构建统一框架融合多种社会交互场景测试智能体
  • 大模型在长程多跳推理和不确定协作中表现明显下降
  • 适合研究多智能体协作、人机协同的学者使用

社会互动中的推理与策略行为是智能的重要标志,远超静态环境下的孤立规划或推理任务。本文提出战略规划、交互与谈判基准(SPIN-Bench),一个跨领域的评估体系,用于衡量智能体在战略规划与社会推理方面的表现。该基准整合经典PDDL任务、竞争性棋类游戏、合作卡牌游戏及多智能体谈判场景,涵盖不同动作空间、状态复杂度与交互智能体数量,模拟多样社会情境。实验表明,当前大模型虽能处理基础事实检索与短期规划,但在涉及大规模状态空间的深度多跳推理及不确定性下的社交协调任务中存在显著性能瓶颈。我们期望SPIN-Bench成为推动鲁棒多智能体规划、社会推理与人-机协同研究的催化剂。

原文摘要 · Abstract (English)

Reasoning and strategic behavior in social interactions is a hallmark of intelligence. This form of reasoning is significantly more sophisticated than isolated planning or reasoning tasks in static settings (e.g., math problem solving). In this paper, we present Strategic Planning, Interaction, and Negotiation (SPIN-Bench), a new multi-domain evaluation designed to measure the intelligence of strategic planning and social reasoning. While many existing benchmarks focus on narrow planning or single-agent reasoning, SPIN-Bench combines classical PDDL tasks, competitive board games, cooperative card games, and multi-agent negotiation scenarios in one unified framework. The framework includes both a benchmark as well as an arena to simulate and evaluate the variety of social settings to test reasoning and strategic behavior of AI agents. We formulate the benchmark SPIN-Bench by systematically varying action spaces, state complexity, and the number of interacting agents to simulate a variety of social settings where success depends on not only methodical and step-wise decision making, but also conceptual inference of other (adversarial or cooperative) participants. Our experiments reveal that while contemporary LLMs handle basic fact retrieval and short-range planning reasonably well, they encounter significant performance bottlenecks in tasks requiring deep multi-hop reasoning over large state spaces and socially adept coordination under uncertainty. We envision SPIN-Bench as a catalyst for future research on robust multi-agent planning, social reasoning, and human--AI teaming. Project Website: https://spinbench.github.io/

多智能体社会推理策略规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。