构建真实投资决策数据集,评估大模型的个性化投资逻辑合理性。
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

- 设计全流程决策追踪框架,包含投资者画像、事件、推理、决策与结果。
- 实测四款大模型逻辑合理率近80%,但事件关联性仅0.8–2.8/5。
- 揭示仅看收益会掩盖推理薄弱问题,适合金融AI研发者参考。
投资能力具有高度个性化:相同市场信息对不同目标、期限、持仓和风险偏好的投资者可能引发不同行动。然而现有金融大模型评估要么依赖静态问答,要么仅看最终盈亏,前者忽略主体性,后者无法判断盈利是基于合理逻辑、符合个人画像,还是偶然运气。我们质疑当前评价体系是否适配高影响决策类智能体。为此提出 extsc{InvestLogicBench},一个面向真实世界的投资决策过程基准,包含151位投资者的201,247条完整决策记录。每条记录以 extbf{P→E→R→D→O} 流程形式呈现:投资者 extit{Profile}、可观测市场 extit{Events}、投资 extit{Reasoning}、可执行 extit{Decision} 及延迟 extit{Outcome}。数据集提供画像构建、时间点事件绑定、结构化逻辑、投资周期、结果与事后复盘,支持理解、画像条件生成与端到端回放。在四款主流大模型上测试显示,逻辑合理性约4/5,但事件关联性仅为0.8–2.8/5;收益表现与过程质量显著脱节。结果揭示了表面流畅却根基薄弱的推理现象,而单纯看回报会掩盖此类问题。我们进一步主张, extbf{P→E→R→D→O} 应成为个性化、高影响智能体的数据系统接口,需支持版本化画像、时间溯源、可检视检索、决策账本与可回放结果。金融场景为这一范式提供了关键验证场,也适用于更广泛的个性化智能代理。
原文摘要 · Abstract (English)
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。