测试大模型代理在不确定企业环境中的长期资源分配能力
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
- 构建132个月的企业财务模拟器,评估长期资源管理
- 仅15.4%的试验完成全程,大模型不必然更优
- 适合研究智能体长期决策与金融自动化的人看
大型语言模型(LLM)代理在复杂任务中表现日益突出,但其在长期视角下对稀缺资源的分配能力仍不明确。与即时反馈的反应性任务不同,此类场景要求代理在部分可观测、后果延迟、资源预算严格和动态变化的条件下做出不可逆承诺。我们引入EnterpriseArena——一个基于真实企业财务数据、匿名业务文档、十年级宏观经济与行业信号及专家验证运营规则构建的132个月金融科技贷款公司首席财务官模拟器。代理需管理流动性、结账、获取高成本信号,并在不同宏观经济环境下申请股权或债务融资。在23个LLM和4种代理框架上的实验表明,当前代理远未达到稳健水平:仅15.4%的试验成功完成整个周期;更大模型并未可靠优于小模型;失败在观测、行动时机与资本规模上呈连锁扩散。这些发现确立了在不确定性下进行长期资源分配作为LLM代理的独特能力缺口。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks with immediate feedback, this setting requires agents to make binding commitments under partial observability, delayed consequences, hard resource budgets, and shifting dynamics. We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes. The simulator is built from transformed firm-level financial data, anonymized business documents, decade-scale macroeconomic and industry signals, and expert-validated operating rules. Experiments across 23 LLMs and four agent frameworks show that current agents remain far from robust: only 15.4% of trials survive the full horizon, larger models do not reliably outperform smaller ones, and failures cascade across observation, action timing, and capital sizing. These findings establish long-horizon resource allocation under uncertainty as a distinct capability gap for LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。