评测金融分析师在真实任务中的推理能力,模型表现远低于人类。
Hedge-Bench: Benchmarking Agents on Hard, Realistic Tasks Pertaining to Financial Reasoning

- 基于真实分析师的工作流程构建102个任务
- 模型平均得分不足16%,差距明显
- 适合评估智能体在复杂金融决策中的能力
AI智能体已能处理金融分析中的机械性任务,如文档检索、公式计算和表格更新。但更具价值的挑战在于对开放式问题进行推理,这正是专业分析师的核心工作。现有基准无法有效捕捉此类问题,且依赖模型自评输出,引入噪声与循环评估。我们提出Hedge-Bench 1.0:包含102个真实职场任务的基准,基于对冲基金分析师的实际工作记录与可验证的推理路径,支持确定性评分。前沿模型与智能体在此基准上平均得分低于16%。数据集与评估工具已开源,地址为github.com/Trata-Inc/trata-hedge-bench。
原文摘要 · Abstract (English)
AI agents can increasingly handle the mechanical tasks of financial analysis: retrieving documents, calculating formulas, updating spreadsheets. The harder, more valuable challenge is reasoning through the open-ended questions that define expert Analyst work. Existing benchmarks do not capture this class of problem, and those that attempt to evaluate open-ended reasoning rely on model-judged outputs that introduce noise and circularity. We present Hedge-Bench 1.0: a benchmark of 102 actual, on-the-job tasks grounded in the explicit reasoning traces of professional hedge fund analysts working with relevant information sources. This approach enables deterministic grading against verified expert steps. Frontier models and agents score below 16\% on the benchmark. We publish the dataset and evaluation harness at github.com/Trata-Inc/trata-hedge-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。