arXiv:2607.24889cs.LGcs.AI2026-07被引 1

为金融模型评估设计新基准,不再依赖唯一标准答案。

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

论文配图:GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
图 1 · 摘自论文原文
  • 基于分析师实际做法构建多层评价体系,避免单一参考值偏差
  • 实测显示顶尖代理得分仅53.4,低于多数初级分析师
  • 适合研究金融智能体、模型评估或投资决策系统的开发者

金融模型结合公开披露信息与分析师假设生成预测和估值。尽管部分环节可机械化验证,但预测结果、折现率和目标价格常存在多个合理答案。现有评测多以单一专家参考为准,但我们基于同一公司独立构建的分析师模型发现:在108个对比对(覆盖65家公司)中,中位单参考得分仅为0.33,92.6%的评分低于0.70,且无同代模型对价格估计一致在10%以内。点对点评分会惩罚专业界本就存在的分歧。为此提出GAUGE基准,评估代理构建的估值模型时参考真实分析师实践,而非单一标准答案。该基准包含1,001份经分类的分析师工作簿和196项任务,采用三层观察实践包络、56个可审计维度、八个有效性关卡及确定性结构检查。通过55人已知群体研究、按公司分组交叉拟合与评审稳定性审计验证。在失败感知分数ϕ₀上,资深分析师均值为88.3,初级为66.0,学生为43.2。24个代理在1,011次生成中最高得分为53.4,高于学生平均但低于所有资深及多数初级分析师。代理通过93%机械维度与78%判断维度,舰队中位差距达26分。当前代理在模型构建上强于估值判断。我们发布方法论、去标识数据层、受控训练划分、版本化48任务核心评估集及预留更新池。

原文摘要 · Abstract (English)

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $ϕ_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

金融模型评估基准智能体评测估值判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。