arXiv:2606.06960cs.CL2026-06被引 2

构建金融领域低重复任务自进化评估基准,测试模型如何从延迟反馈中学习。

FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards

论文配图:FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards
图 1 · 摘自论文原文
  • 基于31个A股行业指数和17.7万篇新闻,构建动态金融信息流
  • 在10/20/40天后评估市场情绪因子预测效果,验证长期反馈有效性
  • 揭示记忆系统在噪声反馈下表现不稳定,适合研究自进化机制的学者

基于自进化语言模型的代理可通过积累和更新测试时经验来改进行为,但现有评估多依赖重复性任务和明确成功信号。本文提出 extsc{FinEvolveBench},一个面向低重复性任务与隐式奖励的自进化代理评估基准。该基准重建了31个中国A股行业指数的日度金融信息流,并将177,324篇公开新闻与市场观测对齐。研究者可设定预测时间窗;本文评估预测市场情绪因子在10、20和40个交易日后的滞后市场调整回报。不同于静态基准独立评分, extsc{FinEvolveBench} 将新决策与早期结果交织,检验代理能否在测试时将嘈杂的现实反馈转化为可复用经验。两个骨干模型实验表明,评估的通用记忆系统未持续优于无经验流水线。匹配消融实验进一步显示,反馈驱动的效用更新在一种模型上短期有效,但在另一种模型上多数情况下反而损害性能。这些结果使 extsc{FinEvolveBench} 成为在噪声、延迟和结果级反馈下诊断基于经验自进化机制的有效测试平台。

原文摘要 · Abstract (English)

Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task patterns and explicit success signals. We introduce \textsc{FinEvolveBench}, a benchmark for self-evolving agents on low-repetition tasks with implicit rewards. The benchmark reconstructs a daily financial information stream over 31 Chinese A-share industry indices and aligns 177,324 public news articles with market observations. Researchers can define prediction horizons over this stream; we evaluate predictive market-sentiment factors against delayed market-adjusted returns after 10, 20, and 40 trading days. Unlike static benchmarks that score each prediction independently, \textsc{FinEvolveBench} interleaves new decisions with delayed outcomes from earlier ones, testing whether agents can convert noisy real-world feedback into reusable experience at test time. Experiments with two backbone models show that the evaluated general-purpose memory systems do not consistently outperform the no-experience pipeline. A matched ablation further shows that feedback-driven utility updates help at shorter horizons on one backbone but hurt in most settings on the other. Together, these results position \textsc{FinEvolveBench} as a diagnostic testbed for experience-based self-evolution under noisy, delayed, and outcome-level feedback.

自进化金融预测延迟反馈评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。