arXiv:2608.06144cs.AI2026-08被引 1

评测金融领域智能体的持续进化能力,看它能否从经验中学习并提升表现。

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

论文配图:FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
图 1 · 摘自论文原文
  • 设计120个真实金融场景任务,分20个业务场景覆盖六大金融领域。
  • 自进化智能体得分提升9.33至19.37点,合规问题减少0.12至0.44个/任务。
  • 仅靠技能进化优于记忆或混合方式,且评分反馈比参考答案更有效。

现有智能体评测难以衡量任务间的经验迁移,且缺乏对专业工作流、开放性交付成果及多维度评估的覆盖。我们提出FinEvo-Bench,一个纵向基准,包含120个基于真实案例的任务,覆盖20个业务场景、六个金融领域。任务依据机构提供的专业流程定义操作与约束,事实来自机构提供或公开文档的真实案例。每个场景含六个相关但实质不同的案例,共享同一流程和人工评审的评分标准。使用相同Qwen3.7-Max模型,对比四种自进化代理框架,并在三个独立随机打乱的全局交错任务流上测试。通过非进化对照组评估进化收益,由Claude Code(基于Claude Opus 4.6)独立评估输出质量。结果表明:Letta获得最高演化得分(91.65),合规问题最少(0.09个/任务);Codex实现最大进化增益(+19.37)。所有框架中,进化条件使得分提升9.33–19.37,合规问题减少0.12–0.44个/任务。场景内后三阶段(秩4-6)的得分提升高于前三个阶段(秩1-3)6.10–8.70点。在独立评测中,仅技能进化优于仅记忆或记忆-技能联合进化。同时,基于评分反馈的优化效果优于参考答案反馈。

原文摘要 · Abstract (English)

Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold's self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.

智能体金融AI自进化评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。