构建可验证的多任务金融智能体评估环境,打破单一任务评价局限
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

- 统一框架整合预测、交易、风控等多阶段金融任务
- 支持论文自动转为可执行任务包,提升评估可复现性
- 含防泄露机制和低延迟回测引擎,适合量化研究与工程落地
尽管大型语言模型代理在量化金融工作流中日益应用,其评估仍分散于孤立任务,且基准任务的金融相关性常被忽视。金融工作流本质上是多阶段的,包含相互依赖的任务如预测、策略构建、风险管理与交易。现有平台通常聚焦单一任务,因而可能夸大代理能力,无法揭示其泛化性、真实市场交互及有意义决策的缺陷。我们提出 OpenFinGym,一个面向量化金融代理开发的统一 Gym 环境,涵盖预测、市场生成、实时交易和欺诈检测,通过统一执行与验证接口实现全流程管理。该环境提供自动化任务构建流水线,可将量化金融论文转化为可执行任务包;具备容器化运行时与宿主侧验证服务,支持规模化代理部署并防止训练-测试数据泄露;配备低延迟数据流设计的纸面交易引擎;支持长期视域与事件驱动型预测的延迟解析;并集成 SFT 与 RL 后训练能力。
原文摘要 · Abstract (English)
Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。