arXiv:2606.17459cs.AI2026-06被引 1

测试大模型当CEO能否合理分配资源,发现其战略判断力仍有限。

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

论文配图:Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
图 1 · 摘自论文原文
  • 让大模型扮演CEO,协调四类高管的矛盾建议做资源分配决策。
  • 所有模型结构有效但战略校准能力差异大,最难点是权衡冲突与果断行动。
  • 适合研究AI辅助决策、组织智能或大模型局限性的读者参考。

评估大语言模型(LLMs)的决策能力日益成为研究重点,但现有基准多聚焦于孤立的认知任务,如推理、知识检索和经济理性,在简化场景中表现良好。然而,真实高管决策的核心挑战在于:在信息不对称、组织约束和时间依赖条件下,整合来自专业化利益相关方的矛盾建议。为此,我们提出 extsc{CEO-Bench},一个基于多智能体的基准,用于评估 LLMs 在企业级战略资源再分配中的表现——即在多轮、高约束的组织环境中,将资本重新分配至不同业务单元。在该基准中,LLM 作为 CEO,接收来自四位角色特化的高管顾问(CFO、CTO、COO、CMO)的冲突建议,每位顾问拥有私有信号和不同优先级,需将其整合为具体分配方案,并从四个维度评估:角色融合度、条件决断力、历史敏感性判断与计划有效性。在五个前沿模型上对13个场景的实验表明,所有模型均具备高结构有效性,但在战略校准层面出现显著分歧。我们识别出系统性失败模式,包括单一顾问主导、模糊情境下的保守默认以及历史遗忘;并发现一种结构性的整合-决断权衡:更深入处理矛盾视角的模型往往产生更不果断的行动。这些发现划定了当前 LLM 作为组织决策者的能力边界,并为未来人工智能辅助高管系统的构建提供指导。

原文摘要 · Abstract (English)

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings. These evaluations overlook the defining challenge of real executive decision-making: integrating conflicting recommendations from specialized stakeholders under information asymmetry, organizational constraints, and temporal dependencies. We introduce \textsc{CEO-Bench}, a multi-agent benchmark that evaluates LLMs on CEO-level strategic resource reallocation -- the process of redirecting capital across business units in a multi-round, constraint-rich organizational environment. In \textsc{CEO-Bench}, LLM agents receive conflicting advice from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO), each with private signals and distinct priorities, and must synthesize these into a concrete allocation plan evaluated along four dimensions: role integration, conditional boldness, history-sensitive judgment, and plan validity. Experiments across five frontier models on 13 scenarios reveal that all models achieve high structural validity but diverge sharply on strategic calibration -- the hardest capability layer. We identify systematic failure modes including single-advisor capture, conservative default under ambiguity, and historical amnesia, and uncover a structural integration-boldness tradeoff: models that engage more deeply with conflicting perspectives tend to produce less decisive action. These findings delineate the current capability boundary of LLMs as organizational decision-makers and inform the design of future AI-assisted executive systems.

大模型决策多智能体战略规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。