测试大模型能否理解并应用投资专家的决策逻辑。
InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

- 构建八层认知评测体系,从原则识别到框架推演
- 发现模型流畅表达但程序推理仍有明显短板
- 适合评估金融AI助手的深层思维能力
大型语言模型日益被用作投资研究助手,但尚无基准测试其是否能准确重构和应用投资专家的具体决策流程。我们提出InvestPhilBench,一个涵盖八个认知层级的多层基准,从原则识别(L1)到新框架外推(L8)。v0.6版本包含118张源自原始资料的原则卡、25张带拓扑元数据的决策框架卡,以及243个问答题(197个开发集/46个保留测试集)。为实现可复现的大规模评分,我们引入基准自动化评分流水线(BASP)、五种算法指标、六类故障模式检测协议(FMDP),以及每门控项的门重建准确率(GRA)。本次发布主要为基准与方法论贡献:其经验研究——在188个问题的开发集上对四个模型进行初步测试(闭卷)——旨在压力测试指标设计而非排名模型。结果表明,不同提供商间存在显著差距(BASP 0.906 vs. 0.438),但这些混合评分是上限估计。核心发现经验证:BASP复合指标在前沿模型(Claude L4=0.932)趋于饱和,而GRA仍暴露程序缺陷(前沿L4 GRA≈0.77,L7 GRA 0.57-0.62)——复合评分奖励流畅表述,掩盖了程序性差距。在100项专家标注的黄金数据集上,自动化BASP复合得分与人工参考的相关系数为0.72(MAE=0.10)。v0.6还实现了统一评分员与真实模型内检索/溯源条件;去混淆的多模型排行榜及完整三条件运行将于v1.0交付。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors. We introduce InvestPhilBench, a multi-layer benchmark spanning eight cognitive tiers, from principle identification (L1) to novel framework extrapolation (L8). The v0.6 release comprises 118 primary-source-verified principle cards, 25 decision-framework cards with explicit topology metadata, and 243 QA questions (197 dev / 46 held-out test). For reproducible scoring at scale we introduce the Benchmark Automated Scoring Pipeline (BASP), five algorithmic metrics, the Failure Mode Detection Protocol (FMDP) covering six failure modes, and Gate Reconstruction Accuracy (GRA), a per-gate metric for questions with gold reasoning programs. This release is primarily a benchmark-and-methodology contribution: its empirical study -- a four-model sanity wave on the 188-question development split (closed-book) -- is deliberately preliminary and stress-tests the metric design rather than ranking models. The wave shows a sharp provider-tier split (BASP 0.906 vs. 0.438), though these mixed-judge numbers are confounded upper bounds. The central methodological finding survives the caveat: the BASP composite saturates at the frontier (Claude L4 = 0.932) while GRA still exposes a procedural deficit (frontier L4 GRA ~0.77, L7 GRA 0.57-0.62) -- composite scoring rewards fluent prose and hides the procedural gap. On a 100-item expert-annotated gold set, the automated BASP composite tracks the human reference at Pearson r = 0.72 (MAE = 0.10). v0.6 also implements a unified judge and true model-in-the-loop retrieval/oracle conditions; the de-confounded multi-model leaderboard and full three-condition run are v1.0 deliverables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。