arXiv:2608.09988cs.CEcs.CL2026-08

提出可审计的LLM投资代理评估框架,杜绝数据泄露与虚假收益。

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

  • 基于五分钟频次的S&P 500实盘模拟,确保决策时可见数据真实可用。
  • 模型在单一冻结窗口下收益率为上限,实际交易成本主要来自换手率。
  • 引入可复现的分析证据与约束强制机制,适合研究者验证代理性能。

大型语言模型越来越多地用于市场解读、风险评估与资本配置。然而,现有LLM交易代理的报告结果可能因前瞻泄露、乐观执行及未执行的风险指令而被夸大。本文提出OpenPM,一个可审计的点时评估框架,用于评估LLM投资代理。在OpenPM中,代理管理100万美元的纯多头投资组合,以五分钟间隔使用S&P 500市场数据。所有对代理可见的数据必须在决策时刻真实可用。自然语言风险指令被转化为类型化约束,并在实际持仓中强制执行。每次运行生成审计材料,包括污染证书、成本敏感性曲线和约束遵守报告。我们还构建了一个基准代理‘分层分配器’:类型化分析师评分标的,构造型LLM生成权重,确定性评议员保证可行性。通过仅捕获一次分析师证据并在不同构造模型间重放,隔离了构造器行为。在短窗口案例研究中,更强的构造器在相同标的池上带来适度且依赖模型的超额收益,但分析师质量比构造器选择更重要,换手率是主要成本来源。所有回报均为单个冻结窗口下的理论上限,未考虑市场冲击,非有效阿尔法。

原文摘要 · Abstract (English)

Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that are described but not enforced. We present OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents. In OpenPM, an agent manages a \$1M long-only book over the S\&P 500 universe using market data at five-minute intervals. Every record visible to the agent must be available at the decision time. Natural-language risk mandates are converted into typed constraints and enforced on the executed portfolio. Each run produces audit artifacts, including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. We also build a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility. We isolate constructor behavior by capturing analyst evidence once and replaying it across constructor models. In our short-window case study, stronger constructors show modest and model-dependent gains over equal weighting on the same pool, but analyst quality matters more than constructor choice, and turnover is the main cost driver. All returns are upper bounds on a single frozen window without market impact, not validated alpha.

LLM投资可审计评估风险控制量化交易

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。