构建金融量化任务评估基准,测试大模型的推理与策略编码能力。
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
- 设计三维度评测体系:知识问答、数学推理、策略代码生成。
- 引入真实回测框架,量化评估模型策略的收益与风险表现。
- 发现大模型在策略编码上远逊于人类专家,但微调后可显著提升。
大型语言模型在多个领域展现强大能力,但在金融量化任务上的评估仍碎片化,主要局限于知识型问答。我们提出QuantEval,一个涵盖三大核心维度的金融量化任务评测基准:基于知识的问答、定量数学推理和定量策略编码。与以往金融评测不同,QuantEval集成类CTA回测框架,可执行模型生成的交易策略并用金融绩效指标评估,实现对量化编码能力的真实测评。我们评估了若干主流开源与专有大模型,发现其在推理与策略编码方面与人类专家存在显著差距。通过大规模监督微调与强化学习实验,在领域对齐数据上实现了持续性能提升。我们希望QuantEval能推动大模型在金融量化领域的研究,并加速其在实际交易流程中的应用。同时,我们公开完整的确定性回测配置(资产池、成本模型、指标定义),确保结果严格可复现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowledge-centric question answering. We introduce QuantEval, a benchmark that evaluates LLMs across three essential dimensions of quantitative finance: knowledge-based QA, quantitative mathematical reasoning, and quantitative strategy coding. Unlike prior financial benchmarks, QuantEval integrates a CTA-style backtesting framework that executes model-generated strategies and evaluates them using financial performance metrics, enabling a more realistic assessment of quantitative coding ability. We evaluate some state-of-the-art open-source and proprietary LLMs and observe substantial gaps to human experts, particularly in reasoning and strategy coding. Finally, we conduct large-scale supervised fine-tuning and reinforcement learning experiments on domain-aligned data, demonstrating consistent improvements. We hope QuantEval will facilitate research on LLMs' quantitative finance capabilities and accelerate their practical adoption in real-world trading workflows. We additionally release the full deterministic backtesting configuration (asset universe, cost model, and metric definitions) to ensure strict reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。