首个金融推理模型Fin-o1通过强化学习提升金融决策能力,效果优于GPT-o1等主流模型。
Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance
- 构建首个高保真金融思维链语料库FinCoT,融合领域监督与迭代优化
- 基于该数据训练的Fin-o1模型在金融推理任务上超越GPT-o1和DeepSeek-R1
- 发现GRPO比PPO、DPO更有效,强调针对性训练优于单纯扩大规模
金融决策中的推理能力对大模型提出独特挑战。尽管强化学习(RL)提升了通用推理能力,但金融领域的进展受限于缺乏有效的金融思维链(CoT)语料库、不同RL方法的系统性对比以及全面的基准测试。为此,我们提出FinCoT,首个开源的高保真金融思维链语料库,通过三阶段管道从七个QA数据集构建,包含领域监督、迭代大模型精炼和难度感知过滤。基于FinCoT,我们开发了首个开源的金融推理模型Fin-o1,采用监督微调与GRPO-based RL训练。实验表明,我们的模型在金融推理任务上优于现有金融模型及主流通用模型如GPT-o1、DeepSeek-R1和GPT-4.5。我们还首次系统评估三种不同RL方法在提升领域推理中的有效性。最后,我们提出FinReason,首个涵盖多表分析、长上下文推理和方程任务的金融推理基准,评估29个LLM。结果显示,通用模型在标准基准表现优异,但在金融场景中性能显著下降;即使经过金融微调的模型如Dianjin-R1和FinR1,在长文档任务中也出现退化。而我们的Fin-o1模型始终优于其基线及更大的GPT-o1和DeepSeek-R1,验证了数据构建与训练策略的有效性。研究进一步表明,GRPO带来稳定增益,而PPO和DPO无明显提升,凸显针对性数据与优化的重要性,而非单纯依赖规模。
原文摘要 · Abstract (English)
As the fundamental capability behind decision-making in finance, financial reasoning poses distinct challenges for LLMs. Although reinforcement learning (RL) have boosted generic reasoning, the progress in finance is hindered by the absence of empirical study of building effective financial chain-of-thought (CoT) corpus, a systematic comparison of different RL methods, and comprehensive benchmarks. To address these gaps, we introduce FinCoT, the first open high-fidelity CoT corpus for finance, distilled from seven QA datasets by a novel three-stage pipeline that incorporates domain supervision, iterative LLM refinement, and difficulty-aware filtering. Based on FinCoT, we develop Fin-o1, the first open financial reasoning models trained via supervised fine-tuning and GRPO-based RL. Our models outperform existing financial reasoning models and SOTA general models such as GPT-o1, DeepSeek-R1, and GPT-4.5. We also investigate the effectiveness of three different RL methods in improving domain-specific reasoning, offering the first such empirical study. We finally propose FinReason, the first financial reasoning benchmark covering multi-table analysis, long-context reasoning, and equation-based tasks, and evaluate 29 LLMs. Our extensive experiments reveal general reasoning models excel on standard benchmarks yet exhibit obvious performance degradation in financial contexts; even finance-tuned models like Dianjin-R1 and FinR1 degrade on lengthy documents. In contrast, our Fin-o1 models consistently outperform their backbones and larger GPT-o1 and DeepSeek-R1, confirming the effectiveness of our data building and model training strategy. Our study further shows that GRPO yields reliable gains whereas PPO and DPO do not, highlighting the need for targeted data and optimisation rather than scale alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。