首个覆盖六类资产的LLM投资组合评估基准,揭示模型推理错误会累积放大。
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

- 构建跨六类资产、2015-2025年全链条投资决策流水线
- 仅32.5%的LLM表现优于等权基准,且推理错误会逐级放大
- 支持实时评估,避免历史数据污染,适合金融AI研究者使用
大语言模型(LLMs)在多种金融任务中表现出色,但投资组合管理(PM)仍缺乏有效评估。现有基准存在两大缺陷:多局限于股票且忽略跨资产相关性;无法评估完整的决策流程。我们提出PortBench,涵盖2015至2025年间六类异构资产。该基准包含6,269个问题的静态QA数据集(七种任务模板)和动态五阶段配置流水线。为评估各环节,引入双层相关性评分(衡量跨类对冲与类内集中度)及CEPS指标,量化推理错误在流程中累积效应。我们在三个压力窗口和三种风险偏好下进行评估,并支持实时评估以缓解历史市场预训练污染。在十种前沿LLM中,强金融问答能力未能转化为优异投资表现:120次评估中仅32.5%在四个市场周期内超越等权基准的夏普比率。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025. PortBench comprises a static QA dataset of 6,269 questions across seven task templates and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress windows and three risk profiles, and support real-time evaluation to mitigate pretraining contamination on historical markets. Across ten frontier LLMs, strong financial QA performance fails to translate into superior portfolio performance: only 32.5\% of 120 evaluations beat equal weighting on Sharpe across four market periods. Our source code is available at \href{https://github.com/AgenticFinLab/portbench}{this https URL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。