arXiv:2605.27896cs.CLcs.CE2026-05

用桌游评测大模型动态理财能力,发现其常因急购资产而陷危机。

FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations

论文配图:FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
图 1 · 摘自论文原文
  • 基于三款金融桌游构建评估框架,测试动态财富管理能力。
  • 9个先进模型虽有长期规划意识,却难有效应对复杂互动获利。
  • 模型易忽视流动性,随机事件下更易陷入财务危机,适合决策系统研究者参考。

近期大语言模型(LLMs)在静态金融推理和简单动态交易任务中表现优异。然而,现有静态金融基准无法充分评估大模型在真实环境中的动态财富管理与金融决策能力。为填补这一空白,我们提出FinBoardBench,一个基于三款经典金融桌游(Cashflow、Acquire、Monopoly)的评估套件。该套件涵盖个人现金流管理与债务平衡、企业投资与收购预测、以及资产拍卖中的竞争性贸易谈判等综合金融技能。对9个先进大模型的实验表明,尽管它们展现出基本的长期规划与投资逻辑,却未能有效利用复杂交互实现盈利;其强大的静态推理能力未转化为成功的动态决策。尤为显著的是,模型倾向于优先获取即时资产,忽视充足流动性,使其在随机事件引发的金融危机中极为脆弱。我们希望FinBoardBench能为未来更智能的基于大模型的决策系统提供重要参考。

原文摘要 · Abstract (English)

Recently, large language models (LLMs) have achieved superior performance in static financial reasoning and simple dynamic trading tasks. However, existing static financial benchmarks are insufficient to assess the dynamic wealth management and financial decision-making capabilities of LLMs in real-world environments. To bridge this gap, we present FinBoardBench, an evaluation suite based on three classic financial board games: Cashflow, Acquire, and Monopoly. FinBoardBench assesses a comprehensive set of financial skills, including personal cash flow management with debt balancing, corporate investment and acquisition forecasting, and competitive trade negotiations with asset auctions. Our experiments with 9 advanced LLMs reveal that while exhibiting basic long-term planning and investment logic, they fail to effectively leverage complex interactions for profit, and their strong static reasoning performance does not transform into successful dynamic decision-making. Notably, they tend to prioritize immediate asset acquisition over maintaining sufficient liquidity, making them vulnerable to financial crises triggered by random events. We hope that FinBoardBench can provide a valuable reference for more intelligent LLM-based decision-making systems in the future.

金融推理大模型评测动态决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。